Terminal-Bench 0.1Benchmark scores and sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: coding · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 GPT-6 Astra OpenAI 64.6%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 0.1Published Sep 3, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
64.6 100% confidence 100 percent, Full
2 Claude Opus 5.5 Anthropic 58.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; production safeguards with fallback; 0.1Published Sep 22, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
58.7 100% confidence 100 percent, Full
3 Claude Fable 5.1 Anthropic 52.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] production safeguards with fallback; 0.1Published Sep 1, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
52.6 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the coding pillar of the SI Score.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed