Terminal-Bench 2.0Benchmark scores and sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: coding · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 GPT-5.5 OpenAI 82.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.0Published Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
82.7 100% confidence 100 percent, Full
2 GPT-5.4 OpenAI 75.1%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.0Published Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
75.1 100% confidence 100 percent, Full
3 Qwen3.7 Max Alibaba / Qwen 69.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.0Published May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
69.7 53% confidence 53 percent, Medium
4 DeepSeek V4 Pro DeepSeek 67.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; 2.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
67.9 85% confidence 85 percent, High
5 MiMo-V2.5 Xiaomi 65.8%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
65.8 88% confidence 88 percent, High
6 GPT-5.4 mini OpenAI 60%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning effort xhigh; 2.0Published Mar 17, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
60.0 100% confidence 100 percent, Full
7 DeepSeek V4 Flash DeepSeek 56.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; 2.0 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
56.9 53% confidence 53 percent, Medium
8 GPT-5.4 nano OpenAI 46.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning effort xhigh; 2.0Published Mar 17, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
46.3 100% confidence 100 percent, Full
9 Laguna XS 2.1 Poolside 37.5%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Harbor; 2.0Published Jul 2, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
37.5 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed