Best AI Models for Reasoning
The strongest general-reasoning models, ranked by the reasoning pillar of the SI Score (GPQA Diamond, Humanity's Last Exam, ARC-AGI-2).
How this list is ranked. Ranked by the reasoning pillar, among models with at least three reasoning results, using the available published results and the current snapshot’s benchmark weights.
The 27 benchmarks in this pillar
- ARC-AGI-1
- ARC-AGI-2
- ARC-AGI-3
- ARC-AGI-1 (public eval)
- ARC-AGI-1 (semi-private)
- ARC-AGI-2 (public eval)
- ARC-AGI-2 (semi-private)
- ARC-AGI-3 (semi-private)
- GPQA Diamond
- Humanity's Last Exam
- Humanity's Last Exam (1,811 verified items)
- Humanity's Last Exam (full set)
- Humanity's Last Exam (full set, text + multimodal)
- Humanity's Last Exam (full set, with tools)
- Humanity's Last Exam (Scale AI)
- Humanity's Last Exam (text only)
- Humanity's Last Exam (text-only subset)
- Humanity's Last Exam (text-only subset, with tools)
- Humanity's Last Exam (text only, with tools)
- Humanity's Last Exam (with tools)
- LiveBench Reasoning: connections
- LiveBench Reasoning: consecutive events
- LiveBench Reasoning: logic with navigation
- LiveBench Reasoning: spatial
- LiveBench Reasoning: theory of mind
- LiveBench Reasoning: zebra puzzles
- MMLU-Pro
- 1 Claude Opus 5.5 Anthropic 91.4
- 2 Muse Spark 1.3 Metaprovisional 90.5
- 3 Muse Spark 1.2 Metaprovisional 90.4
- 4 Claude Fable 5 Anthropicprovisional 89.4
- 5 Claude Sonnet 5.5 Anthropic 88.6
- 6 GPT-6.1 Sol OpenAI 87.5
- 7 Claude Sonnet 5 Anthropicprovisional 87.3
- 8 Claude Fable 5.1 Anthropic 87.1
- 9 GPT-6 Astra OpenAI 87.1
- 10 GPT-5.5 Pro OpenAI 85.9
- 11 Qwen3.8 Max Alibaba / Qwenprovisional 85.5
- 12 Muse Spark 1.1 Metaprovisional 85.3
- 13 GLM-5.3 Z.aiprovisional 84.4
- 14 Claude Opus 5 Anthropicprovisional 84.3
- 15 Gemini 3.7 Flash Google 83.9
Updated Oct 9, 2026. Task lists rank on one stated measure rather than the blended SI Score. See the methodology for how pillars and confidence are computed; a dash means the source has not reported that value.