Best AI Models for Coding

Coding and agentic models ordered by their published coding pillar, combining the benchmark variants available in this snapshot.

How this list is ranked. Ranked by the coding pillar, among models with at least three coding results. Its benchmark definitions and within-pillar weights come from the current snapshot below; model pages show the harness, variant and source for each result.

The 17 benchmarks in this pillar
  1. 1 DeepSeek V4.1 Flash DeepSeek 73.3 66.2 $0.15 1M
  2. 2 Claude Opus 5.5 Anthropic 72.5 78.4 $4.00 1M
  3. 3 Kimi K3 Moonshot AIprovisional 71.1 71.0 $3.00 1M
  4. 4 Claude Fable 5 Anthropicprovisional 70.2 76.8 $10.00 1M
  5. 5 Claude Fable 5.1 Anthropic 69.8 80.2 $10.00 1M
  6. 6 Claude Opus 5 Anthropicprovisional 69.7 74.9 $5.00 1M
  7. 7 Hy3 Tencentprovisional 69.2 60.7 $0.00 256K
  8. 8 Qwen3.8 Max Alibaba / Qwenprovisional 69.1 68.7 $1.65 1M
  9. 9 Qwen3.8 27B Alibaba / Qwenprovisional 67.1 61.9 $0.50 262K
  10. 10 GPT-5 OpenAI 66.7 55.7 $1.25 400K
  11. 11 Muse Spark 1.2 Metaprovisional 66.7 62.3 $1.25 1M
  12. 12 GPT-5.5 OpenAI 66.6 69.5 $5.00 1.1M
  13. 13 Claude Opus 4.7 Anthropicprovisional 66.5 71.2 $5.00 1M
  14. 14 GLM-5.2 Z.aiprovisional 66.2 65.7 $1.40 1M
  15. 15 DeepSeek V4 Pro 0813 DeepSeekprovisional 65.4 63.0 $0.66 1M

Updated Oct 9, 2026. Task lists rank on one stated measure rather than the blended SI Score. See the methodology for how pillars and confidence are computed; a dash means the source has not reported that value.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed