Terminal-Bench 2.1Benchmark scores and sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: coding · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 DeepSeek V4.1 Flash DeepSeek 90.6%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; DeepSeek Harness minimal mode; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
90.6 69% confidence 69 percent, Medium
2 DeepSeek V4.1 Flash DeepSeek 90.3%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; mini-swe-agent; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
90.3 69% confidence 69 percent, Medium
3 MiMo-V2.6-Pro Xiaomi 89.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] MiMo-V2.6 Pro comparison column; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
89.9 88% confidence 88 percent, High
4 Gemini 3.8 Flash Google 89.4%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus 2; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
89.4 100% confidence 100 percent, Full
5 Muse Spark 1.3 Meta 88.8%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Muse Code; 2.1Published Sep 2, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.8 93% confidence 93 percent, High
6 GPT-5.6 Sol OpenAI 88.8%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.8 100% confidence 100 percent, Full
7 Kimi K3 Moonshot AI 88.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Kimi Code; 2.1Published Jul 16, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.3 100% confidence 100 percent, Full
8 GLM-5.3 Z.ai 88.2%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Claude Code 2.1.207; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.2 93% confidence 93 percent, High
9 Claude Fable 5 Anthropic 88%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.0 100% confidence 100 percent, Full
10 Claude Mythos 5 Anthropic 88%Anthropic Mythos 5 system-card claimsLab claim, not an independent evaluation. mini-SWE-agent, GKE, 1x timeout and 3x memory; high effort; 445 trials; June 9 release card, section 8; immutable edition [variant] Mythos 5 system cardPublished Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.0 5% confidence 5 percent, Low
11 DeepSeek V4.1 Flash DeepSeek 88%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; Claude Code; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.0 69% confidence 69 percent, Medium
12 DeepSeek V4 Pro 0813 DeepSeek 87.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; DeepSeek Harness minimal mode; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
87.9 69% confidence 69 percent, Medium
13 MiMo-V2.6-Flash Xiaomi 87.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
87.6 88% confidence 88 percent, High
14 GPT-5.6 Terra OpenAI 87.4%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
87.4 100% confidence 100 percent, Full
15 Qwen3.8 Max Alibaba / Qwen 86.6%Official model cards via models.devLab-reported; metric avg@10; transcribed by MIT models.dev catalog; not independently evaluated [variant] 5h timeout; Claude Code; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
86.6 85% confidence 85 percent, High
16 Qwen3.8 Max Preview Alibaba / Qwen 86.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh; 2.1Published Aug 3, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
86.6 6% confidence 6 percent, Low
17 DeepSeek V4.1 Flash DeepSeek 86.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; Pi; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
86.1 69% confidence 69 percent, Medium
18 DeepSeek V4.1 Flash DeepSeek 85.8%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; DeepSeek Harness PTC mode; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
85.8 69% confidence 69 percent, Medium
19 DeepSeek V4.1 Flash DeepSeek 85.8%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; DeepSeek Harness standard mode; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
85.8 69% confidence 69 percent, Medium
20 Gemini 3.7 Flash Google 85.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Aug 13, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
85.8 100% confidence 100 percent, Full
21 DeepSeek V4.1 Flash DeepSeek 85%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; OpenCode; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
85.0 69% confidence 69 percent, Medium
22 GPT-5.6 Luna OpenAI 84.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
84.7 100% confidence 100 percent, Full
23 GLM-5.3-Flash Z.ai 84.3%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Claude Code 2.1.207; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
84.3 100% confidence 100 percent, Full
24 DeepSeek V4.1 Flash DeepSeek 84.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; Codex; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
84.1 69% confidence 69 percent, Medium
25 DeepSeek V4 Flash Vision Exp DeepSeek 83.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; DeepSeek Harness minimal mode; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
83.9 21% confidence 21 percent, Low
26 Grok 4.5 xAI 83.3%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 8, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
83.3 100% confidence 100 percent, Full
27 Muse Spark 1.2 Meta 82.9%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh effort; Muse Code; 2.1Published Sep 2, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
82.9 53% confidence 53 percent, Medium
28 DeepSeek V4 Flash 0731 DeepSeek 82.7%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] max; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
82.7 69% confidence 69 percent, Medium
29 GLM-5.2 Z.ai 82.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Claude Code; 2.1Published Jun 16, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
82.7 100% confidence 100 percent, Full
30 GLM-5.2 Z.ai 81%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus 2; 2.1Published Jun 16, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
81.0 100% confidence 100 percent, Full
31 Claude Sonnet 5 Anthropic 80.4%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published Jun 30, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
80.4 93% confidence 93 percent, High
32 Muse Spark 1.1 Meta 80%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
80.0 61% confidence 61 percent, Medium
33 GPT-5.5 OpenAI 78.2%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 28, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
78.2 100% confidence 100 percent, Full
34 Gemini 3.6 Flash Google 78%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus 2; 2.1Published Jul 21, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
78.0 100% confidence 100 percent, Full
35 Gemini 3.5 Flash Google 76.2%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 19, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
76.2 100% confidence 100 percent, Full
36 Claude Opus 4.8 Anthropic 74.6%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 28, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
74.6 100% confidence 100 percent, Full
37 Qwen3.8 27B Alibaba / Qwen 73%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
73.0 69% confidence 69 percent, Medium
38 Hy3 Tencent 71.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] highest reasoning effort; 4h timeout; 500 episodes; Terminus 2; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
71.7 88% confidence 88 percent, High
39 LongCat-2.0 Meituan 70.8%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 30, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
70.8 100% confidence 100 percent, Full
40 Gemini 3.1 Pro Preview Google 70.3%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 28, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
70.3 100% confidence 100 percent, Full
41 Claude Sonnet 4.6 Anthropic 67%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published Jun 30, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
67.0 100% confidence 100 percent, Full
42 Claude Opus 4.7 Anthropic 66.1%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 28, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
66.1 100% confidence 100 percent, Full
43 MiniMax-M3 MiniMax 66%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 1, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
66.0 100% confidence 100 percent, Full
44 MiMo-V2.5-Pro Xiaomi 65.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] MiMo-V2.5 Pro comparison column; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
65.2 88% confidence 88 percent, High
45 Step 3.7 Flash StepFun 59.6%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published May 29, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
59.6 13% confidence 13 percent, Low
46 Nemotron 3 Ultra 550B A55B NVIDIA 56.4%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 4, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
56.4 100% confidence 100 percent, Full
47 Gemini 3.5 Flash Lite Google 54%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus 2; 2.1Published Jul 21, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
54.0 100% confidence 100 percent, Full
48 Muse Glimmer 30B Meta 51.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] with terminus2; 2.1Published Aug 10, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
51.7 6% confidence 6 percent, Low
49 MiniMax-M2.7 MiniMax 51.1%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 1, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
51.1 88% confidence 88 percent, High
50 Nemotron 3.5 Lightning 30B A3B NVIDIA 24.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; NeMo Evaluator; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
24.6 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed