Best AI Models for Coding
Coding and agentic models ordered by their published coding pillar, combining the benchmark variants available in this snapshot.
How this list is ranked. Ranked by the coding pillar, among models with at least three coding results. Its benchmark definitions and within-pillar weights come from the current snapshot below; model pages show the harness, variant and source for each result.
The 17 benchmarks in this pillar
- Aider Polyglot
- LiveBench Coding: code completion
- LiveBench Coding: code generation
- LiveBench Coding: JavaScript
- LiveBench Coding: Python
- LiveBench Coding: TypeScript
- SWE-bench Pro
- SWE-bench Pro (public)
- SWE-bench Pro (Qwen-corrected tasks)
- SWE-bench Pro (system card, Jun 2026)
- SWE-bench Verified
- Terminal-Bench
- Terminal-Bench 0.1
- Terminal-Bench 2.0
- Terminal-Bench 2.1
- Terminal-Bench 3.0
- Terminal-Bench 4.0
- 1 DeepSeek V4.1 Flash DeepSeek 73.3
- 2 Claude Opus 5.5 Anthropic 72.5
- 3 Kimi K3 Moonshot AIprovisional 71.1
- 4 Claude Fable 5 Anthropicprovisional 70.2
- 5 Claude Fable 5.1 Anthropic 69.8
- 6 Claude Opus 5 Anthropicprovisional 69.7
- 7 Hy3 Tencentprovisional 69.2
- 8 Qwen3.8 Max Alibaba / Qwenprovisional 69.1
- 9 Qwen3.8 27B Alibaba / Qwenprovisional 67.1
- 10 GPT-5 OpenAI 66.7
- 11 Muse Spark 1.2 Metaprovisional 66.7
- 12 GPT-5.5 OpenAI 66.6
- 13 Claude Opus 4.7 Anthropicprovisional 66.5
- 14 GLM-5.2 Z.aiprovisional 66.2
- 15 DeepSeek V4 Pro 0813 DeepSeekprovisional 65.4
Updated Oct 9, 2026. Task lists rank on one stated measure rather than the blended SI Score. See the methodology for how pillars and confidence are computed; a dash means the source has not reported that value.