Anthropic Alignment, in collaboration with Anthropic, has introduced the Conceptual Reasoning Index (CRI), a new suite of three benchmarks designed to measure AI models' conceptual reasoning capabilities. This index addresses tasks where empirical feedback loops are absent, requiring argumentation similar to philosophy or AI futurism. The CRI aggregates performance across three distinct benchmarks: LMCA, ACCoRD, and DTBench capabilities.
Key Points
- The Conceptual Reasoning Index (CRI) is an aggregate of three benchmarks: LMCA (60%), ACCoRD (20%), and DTBench capabilities (20%).
- LMCA (Language Model Conceptual Argumentation) contains 560 position texts and 1,461 arguments, with 2,140 expert ratings primarily by Emery Cooper.
- ACCoRD (Assessment of Consistency in Conceptual Reasoning Domains) includes 567 human-approved consistency constraints from nearly 14,000 model-generated constraints across 18 types.
- DTBench capabilities (Decision Theory Benchmark) consists of 407 handcrafted multiple-choice questions focused on decision-theoretic situations, created by Caspar Oesterheld and validated by Emery Cooper.
- The CRI currently measures models' ability to judge arguments, with future plans to include argumentation ability.
- Model scores on the CRI were last updated as of August 10, 2026, including Claude Fable 5, Muse Spark 1.2, and Gemini 3.6 Flash.
Context
According to Anthropic Alignment, AI models typically perform less effectively on tasks that cannot be empirically or mathematically verified due to their reliance on abundant data and reliable performance feedback during training. Reducing risks from advanced AI often involves such tasks, which necessitate strong conceptual reasoning. The development of the CRI aims to measure this specific capability, which is defined as the ability to reason about questions with limited empirical evidence, no practically verifiable answer, and a heavy reliance on argumentation.
LMCA evaluates models' ability to judge arguments against position texts by comparing model ratings to expert ratings based on a detailed rubric. ACCoRD assesses logical consistency in models' reported beliefs and preferences on conceptual issues, using numeric probability estimates or preference orderings. DTBench capabilities measures reasoning in decision-theoretic scenarios involving self-prediction or interactions with copies.
Why It Matters
Improving AI models' conceptual reasoning is presented as a critical factor for managing AI risks, as it could enable AIs to assist in understanding complex situations, planning, and developing mitigations. For builders and researchers, the CRI offers a structured approach to evaluate model performance on tasks that are difficult to assess through traditional empirical means, highlighting areas where models may need selective improvement to address advanced AI risks.
What To Do
- Request access to the LMCA dataset via the provided form to explore its curated conceptual arguments.
- Review the methodology details available at conceptualreasoning.ai to understand the construction of the CRI and its constituent benchmarks.
- Note the specific weighting of LMCA (60%), ACCoRD (20%), and DTBench capabilities (20%) when interpreting overall CRI scores.
- Compare the performance of models like Claude Fable 5, Muse Spark 1.2, and Gemini 3.6 Flash on the CRI against their performance on other benchmarks.