Best AI models by taskCoding, math, reasoning, long context, open weights and value

The leaderboard answers which model is best overall. These lists answer best for what, each ranked on one stated measure, with every number linked to its source.

  1. Coding

    Coding and agentic models ordered by their published coding pillar, combining the benchmark variants available in this snapshot.

    1. 1 DeepSeek V4.1 FlashDeepSeek 73.3
    2. 2 Claude Opus 5.5Anthropic 72.5
    3. 3 Kimi K3Moonshot AI 71.1
    Full list
  2. Math

    Mathematical-reasoning models ordered by the math pillar from the available published evaluations.

    1. 1 GPT-6.1 SolOpenAI 93.0
    2. 2 GPT-6 AstraOpenAI 92.9
    3. 3 Claude Opus 5.5Anthropic 91.5
    Full list
  3. Reasoning

    The strongest general-reasoning models, ranked by the reasoning pillar of the SI Score (GPQA Diamond, Humanity's Last Exam, ARC-AGI-2).

    1. 1 Claude Opus 5.5Anthropic 91.4
    2. 2 Muse Spark 1.3Meta 90.5
    3. 3 Muse Spark 1.2Meta 90.4
    Full list
  4. Long context

    Models ordered by advertised input context window, with their SI Score alongside so you can trade reach against quality.

    1. 1 Llama 4 Scout 17B InstructMeta 10M
    2. 2 GPT-5.4OpenAI 1.1M
    3. 3 GPT-5.4 ProOpenAI 1.1M
    Full list
  5. Open weights

    Open-weight models ranked by SI Score, with the exact license each weight release carries shown next to it.

    1. 1 Kimi K3Moonshot AI 71.0
    2. 2 MiniMax-M3MiniMax 68.0
    3. 3 GLM-5.3Z.ai 67.7
    Full list
  6. Price and performance

    The most SI Score per dollar, using a blended price of input and output tokens at a 3:1 mix.

    1. 1 Gemma 4 26B A4B ITGoogle 253.6
    2. 2 Gemma 4 31B ITGoogle 252.0
    3. 3 Hy3Tencent 242.8
    Full list

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed