Terminal-Bench 2.1Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash DeepSeek | 90.6%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; DeepSeek Harness minimal mode; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 90.6 | 69% confidence 69 percent, Medium |
| 2 | DeepSeek V4.1 Flash DeepSeek | 90.3%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; mini-swe-agent; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 90.3 | 69% confidence 69 percent, Medium |
| 3 | MiMo-V2.6-Pro Xiaomi | 89.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] MiMo-V2.6 Pro comparison column; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 89.9 | 88% confidence 88 percent, High |
| 4 | Gemini 3.8 Flash Google | 89.4%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus 2; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 89.4 | 100% confidence 100 percent, Full |
| 5 | Muse Spark 1.3 Meta | 88.8%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Muse Code; 2.1Published Sep 2, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 88.8 | 93% confidence 93 percent, High |
| 6 | GPT-5.6 Sol OpenAI | 88.8%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 88.8 | 100% confidence 100 percent, Full |
| 7 | Kimi K3 Moonshot AI | 88.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Kimi Code; 2.1Published Jul 16, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 88.3 | 100% confidence 100 percent, Full |
| 8 | GLM-5.3 Z.ai | 88.2%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Claude Code 2.1.207; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 88.2 | 93% confidence 93 percent, High |
| 9 | Claude Fable 5 Anthropic | 88%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 88.0 | 100% confidence 100 percent, Full |
| 10 | Claude Mythos 5 Anthropic | 88%Anthropic Mythos 5 system-card claimsLab claim, not an independent evaluation. mini-SWE-agent, GKE, 1x timeout and 3x memory; high effort; 445 trials; June 9 release card, section 8; immutable edition [variant] Mythos 5 system cardPublished Jun 9, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.0 | 5% confidence 5 percent, Low |
| 11 | DeepSeek V4.1 Flash DeepSeek | 88%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; Claude Code; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 88.0 | 69% confidence 69 percent, Medium |
| 12 | DeepSeek V4 Pro 0813 DeepSeek | 87.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; DeepSeek Harness minimal mode; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 87.9 | 69% confidence 69 percent, Medium |
| 13 | MiMo-V2.6-Flash Xiaomi | 87.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 87.6 | 88% confidence 88 percent, High |
| 14 | GPT-5.6 Terra OpenAI | 87.4%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 87.4 | 100% confidence 100 percent, Full |
| 15 | Qwen3.8 Max Alibaba / Qwen | 86.6%Official model cards via models.devLab-reported; metric avg@10; transcribed by MIT models.dev catalog; not independently evaluated [variant] 5h timeout; Claude Code; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 86.6 | 85% confidence 85 percent, High |
| 16 | Qwen3.8 Max Preview Alibaba / Qwen | 86.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh; 2.1Published Aug 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 86.6 | 6% confidence 6 percent, Low |
| 17 | DeepSeek V4.1 Flash DeepSeek | 86.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; Pi; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 86.1 | 69% confidence 69 percent, Medium |
| 18 | DeepSeek V4.1 Flash DeepSeek | 85.8%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; DeepSeek Harness PTC mode; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 85.8 | 69% confidence 69 percent, Medium |
| 19 | DeepSeek V4.1 Flash DeepSeek | 85.8%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; DeepSeek Harness standard mode; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 85.8 | 69% confidence 69 percent, Medium |
| 20 | Gemini 3.7 Flash Google | 85.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Aug 13, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 85.8 | 100% confidence 100 percent, Full |
| 21 | DeepSeek V4.1 Flash DeepSeek | 85%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; OpenCode; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 85.0 | 69% confidence 69 percent, Medium |
| 22 | GPT-5.6 Luna OpenAI | 84.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 84.7 | 100% confidence 100 percent, Full |
| 23 | GLM-5.3-Flash Z.ai | 84.3%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Claude Code 2.1.207; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 84.3 | 100% confidence 100 percent, Full |
| 24 | DeepSeek V4.1 Flash DeepSeek | 84.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; Codex; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 84.1 | 69% confidence 69 percent, Medium |
| 25 | DeepSeek V4 Flash Vision Exp DeepSeek | 83.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; DeepSeek Harness minimal mode; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 83.9 | 21% confidence 21 percent, Low |
| 26 | Grok 4.5 xAI | 83.3%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 8, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 83.3 | 100% confidence 100 percent, Full |
| 27 | Muse Spark 1.2 Meta | 82.9%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh effort; Muse Code; 2.1Published Sep 2, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 82.9 | 53% confidence 53 percent, Medium |
| 28 | DeepSeek V4 Flash 0731 DeepSeek | 82.7%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] max; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 82.7 | 69% confidence 69 percent, Medium |
| 29 | GLM-5.2 Z.ai | 82.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Claude Code; 2.1Published Jun 16, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 82.7 | 100% confidence 100 percent, Full |
| 30 | GLM-5.2 Z.ai | 81%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus 2; 2.1Published Jun 16, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 81.0 | 100% confidence 100 percent, Full |
| 31 | Claude Sonnet 5 Anthropic | 80.4%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published Jun 30, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 80.4 | 93% confidence 93 percent, High |
| 32 | Muse Spark 1.1 Meta | 80%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jul 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 80.0 | 61% confidence 61 percent, Medium |
| 33 | GPT-5.5 OpenAI | 78.2%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 28, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 78.2 | 100% confidence 100 percent, Full |
| 34 | Gemini 3.6 Flash Google | 78%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus 2; 2.1Published Jul 21, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 78.0 | 100% confidence 100 percent, Full |
| 35 | Gemini 3.5 Flash Google | 76.2%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 19, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 76.2 | 100% confidence 100 percent, Full |
| 36 | Claude Opus 4.8 Anthropic | 74.6%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 28, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 74.6 | 100% confidence 100 percent, Full |
| 37 | Qwen3.8 27B Alibaba / Qwen | 73%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 73.0 | 69% confidence 69 percent, Medium |
| 38 | Hy3 Tencent | 71.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] highest reasoning effort; 4h timeout; 500 episodes; Terminus 2; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 71.7 | 88% confidence 88 percent, High |
| 39 | LongCat-2.0 Meituan | 70.8%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 30, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 70.8 | 100% confidence 100 percent, Full |
| 40 | Gemini 3.1 Pro Preview Google | 70.3%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 28, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 70.3 | 100% confidence 100 percent, Full |
| 41 | Claude Sonnet 4.6 Anthropic | 67%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published Jun 30, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 67.0 | 100% confidence 100 percent, Full |
| 42 | Claude Opus 4.7 Anthropic | 66.1%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus-2; 2.1Published May 28, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 66.1 | 100% confidence 100 percent, Full |
| 43 | MiniMax-M3 MiniMax | 66%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 1, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 66.0 | 100% confidence 100 percent, Full |
| 44 | MiMo-V2.5-Pro Xiaomi | 65.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] MiMo-V2.5 Pro comparison column; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 65.2 | 88% confidence 88 percent, High |
| 45 | Step 3.7 Flash StepFun | 59.6%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published May 29, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 59.6 | 13% confidence 13 percent, Low |
| 46 | Nemotron 3 Ultra 550B A55B NVIDIA | 56.4%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 4, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 56.4 | 100% confidence 100 percent, Full |
| 47 | Gemini 3.5 Flash Lite Google | 54%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Terminus 2; 2.1Published Jul 21, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 54.0 | 100% confidence 100 percent, Full |
| 48 | Muse Glimmer 30B Meta | 51.7%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] with terminus2; 2.1Published Aug 10, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 51.7 | 6% confidence 6 percent, Low |
| 49 | MiniMax-M2.7 MiniMax | 51.1%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 1, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 51.1 | 88% confidence 88 percent, High |
| 50 | Nemotron 3.5 Lightning 30B A3B NVIDIA | 24.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; NeMo Evaluator; 2.1
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 24.6 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.