Best AI Models for Reasoning

The strongest general-reasoning models, ranked by the reasoning pillar of the SI Score (GPQA Diamond, Humanity's Last Exam, ARC-AGI-2).

  1. 1 Claude Opus 5.5 Anthropic 91.4 78.4 $4.00 1M
  2. 2 Muse Spark 1.3 Metaprovisional 90.5 70.1 $1.25 1M
  3. 3 Muse Spark 1.2 Metaprovisional 90.4 62.3 $1.25 1M
  4. 4 Claude Fable 5 Anthropicprovisional 89.4 76.8 $10.00 1M
  5. 5 Claude Sonnet 5.5 Anthropic 88.6 69.4 $2.00 1M
  6. 6 GPT-6.1 Sol OpenAI 87.5 73.0 $2.00 1.1M
  7. 7 Claude Sonnet 5 Anthropicprovisional 87.3 67.7 $2.00 1M
  8. 8 Claude Fable 5.1 Anthropic 87.1 80.2 $10.00 1M
  9. 9 GPT-6 Astra OpenAI 87.1 77.5 $10.00 1.1M
  10. 10 GPT-5.5 Pro OpenAI 85.9 63.5 $30.00 1.1M
  11. 11 Qwen3.8 Max Alibaba / Qwenprovisional 85.5 68.7 $1.65 1M
  12. 12 Muse Spark 1.1 Metaprovisional 85.3 62.9 $1.25 1M
  13. 13 GLM-5.3 Z.aiprovisional 84.4 67.7 $1.40 1M
  14. 14 Claude Opus 5 Anthropicprovisional 84.3 74.9 $5.00 1M
  15. 15 Gemini 3.7 Flash Google 83.9 70.0 $0.75 1M

Updated Oct 9, 2026. Task lists rank on one stated measure rather than the blended SI Score. See the methodology for how pillars and confidence are computed; a dash means the source has not reported that value.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed