Terminal-Bench 3.0Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | DeepSeek V4.1 Flash DeepSeek | 30%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; DeepSeek Harness minimal mode; 3.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 30.0 | 69% confidence 69 percent, Medium |
| 2 | GLM-5.3 Z.ai | 28.3%Official model cards via models.devLab-reported; metric avg@3; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Claude Code 2.1.207; 3.0
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 28.3 | 93% confidence 93 percent, High |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.