Best AI Models for Math

Mathematical-reasoning models ordered by the math pillar from the available published evaluations.

How this list is ranked. Ranked by the math pillar, among models with at least three math results, using the snapshot’s benchmark weights. Check coverage and individual evaluations before treating a small lead as decisive.

The 13 benchmarks in this pillar
  1. 1 GPT-6.1 Sol OpenAI 93.0 73.0 $2.00 1.1M
  2. 2 GPT-6 Astra OpenAI 92.9 77.5 $10.00 1.1M
  3. 3 Claude Opus 5.5 Anthropic 91.5 78.4 $4.00 1M
  4. 4 Claude Fable 5.1 Anthropic 91.5 80.2 $10.00 1M
  5. 5 GPT-6 Sol OpenAI 91.3 68.9 $2.00 1.1M
  6. 6 Claude Fable 5 Anthropicprovisional 91.0 76.8 $10.00 1M
  7. 7 Claude Sonnet 5.5 Anthropic 90.5 69.4 $2.00 1M
  8. 8 GPT-5.6 Sol OpenAIprovisional 89.3 72.1 $4.00 1.1M
  9. 9 Mistral Large 4 Mistral AIprovisional 88.1 62.2 $0.68 1M
  10. 10 DeepSeek V4.1 Flash DeepSeek 88.0 66.2 $0.15 1M
  11. 11 Claude Opus 5 Anthropicprovisional 86.9 74.9 $5.00 1M
  12. 12 Muse Spark 1.2 Metaprovisional 86.7 62.3 $1.25 1M
  13. 13 Nemotron 3 Ultra 550B A55B NVIDIAprovisional 85.5 61.2 — 1M
  14. 14 GPT-5.6 Terra OpenAIprovisional 85.1 68.2 $2.00 1.1M
  15. 15 Muse Spark 1.3 Metaprovisional 83.2 70.1 $1.25 1M

Updated Oct 9, 2026. Task lists rank on one stated measure rather than the blended SI Score. See the methodology for how pillars and confidence are computed; a dash means the source has not reported that value.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed