Best AI models by taskCoding, math, reasoning, long context, open weights and value
The leaderboard answers which model is best overall. These lists answer best for what, each ranked on one stated measure, with every number linked to its source.
-
Coding
Coding and agentic models ordered by their published coding pillar, combining the benchmark variants available in this snapshot.
- 1 DeepSeek V4.1 FlashDeepSeek 73.3
- 2 Claude Opus 5.5Anthropic 72.5
- 3 Kimi K3Moonshot AI 71.1
-
Math
Mathematical-reasoning models ordered by the math pillar from the available published evaluations.
- 1 GPT-6.1 SolOpenAI 93.0
- 2 GPT-6 AstraOpenAI 92.9
- 3 Claude Opus 5.5Anthropic 91.5
-
Reasoning
The strongest general-reasoning models, ranked by the reasoning pillar of the SI Score (GPQA Diamond, Humanity's Last Exam, ARC-AGI-2).
- 1 Claude Opus 5.5Anthropic 91.4
- 2 Muse Spark 1.3Meta 90.5
- 3 Muse Spark 1.2Meta 90.4
-
Long context
Models ordered by advertised input context window, with their SI Score alongside so you can trade reach against quality.
- 1 Llama 4 Scout 17B InstructMeta 10M
- 2 GPT-5.4OpenAI 1.1M
- 3 GPT-5.4 ProOpenAI 1.1M
-
Open weights
Open-weight models ranked by SI Score, with the exact license each weight release carries shown next to it.
- 1 Kimi K3Moonshot AI 71.0
- 2 MiniMax-M3MiniMax 68.0
- 3 GLM-5.3Z.ai 67.7
-
Price and performance
The most SI Score per dollar, using a blended price of input and output tokens at a 3:1 mix.
- 1 Gemma 4 26B A4B ITGoogle 253.6
- 2 Gemma 4 31B ITGoogle 252.0
- 3 Hy3Tencent 242.8