Humanity's Last Exam (with tools)Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Claude Opus 5.5 Anthropic | 67.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; with tools; production safeguards with fallbackPublished Sep 22, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 67.7 | 100% confidence 100 percent, Full |
| 2 | Claude Fable 5.1 Anthropic | 65%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools; production safeguards with fallbackPublished Sep 1, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 65.0 | 100% confidence 100 percent, Full |
| 3 | Claude Opus 5 Anthropic | 64.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jul 24, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 64.7 | 100% confidence 100 percent, Full |
| 4 | Claude Fable 5 Anthropic | 64.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 64.5 | 100% confidence 100 percent, Full |
| 5 | Claude Sonnet 5.5 Anthropic | 64.5%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Sep 28, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 64.5 | 93% confidence 93 percent, High |
| 6 | DeepSeek V4.1 Flash DeepSeek | 63.9%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; with tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 63.9 | 69% confidence 69 percent, Medium |
| 7 | GLM-5.3 Z.ai | 62.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; with tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 62.5 | 93% confidence 93 percent, High |
| 8 | Muse Spark 1.1 Meta | 62.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jul 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 62.1 | 61% confidence 61 percent, Medium |
| 9 | DeepSeek V4 Pro 0813 DeepSeek | 60%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 60.0 | 69% confidence 69 percent, Medium |
| 10 | GPT-5.4 Pro OpenAI | 58.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 58.7 | 69% confidence 69 percent, Medium |
| 11 | Claude Opus 4.8 Anthropic | 57.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 57.9 | 100% confidence 100 percent, Full |
| 12 | Claude Haiku 5.5 Anthropic | 57.4%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 57.4 | 68% confidence 68 percent, Medium |
| 13 | GPT-5.5 Pro OpenAI | 57.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 57.2 | 53% confidence 53 percent, Medium |
| 14 | GPT-6 Astra OpenAI | 57.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Sep 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 57.2 | 100% confidence 100 percent, Full |
| 15 | Qwen3.8 Max Alibaba / Qwen | 56.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 56.2 | 85% confidence 85 percent, High |
| 16 | Qwen3.8 Max Preview Alibaba / Qwen | 56.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh, with toolsPublished Aug 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 56.2 | 6% confidence 6 percent, Low |
| 17 | Claude Opus 4.7 Anthropic | 54.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 54.7 | 100% confidence 100 percent, Full |
| 18 | GPT-5.5 OpenAI | 52.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 52.2 | 100% confidence 100 percent, Full |
| 19 | GPT-5.4 OpenAI | 52.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 52.1 | 100% confidence 100 percent, Full |
| 20 | DeepSeek V4 Pro DeepSeek | 48.2%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; with tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 48.2 | 85% confidence 85 percent, High |
| 21 | Step 3.7 Flash StepFun | 47.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished May 29, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 47.2 | 13% confidence 13 percent, Low |
| 22 | Claude Sonnet 4.6 Anthropic | 46.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 30, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 46.8 | 100% confidence 100 percent, Full |
| 23 | DeepSeek V4 Flash DeepSeek | 45.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; with tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 45.1 | 53% confidence 53 percent, Medium |
| 24 | GPT-5.4 mini OpenAI | 41.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Mar 17, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 41.5 | 100% confidence 100 percent, Full |
| 25 | GPT-5.4 nano OpenAI | 37.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Mar 17, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 37.7 | 100% confidence 100 percent, Full |
| 26 | Nemotron 3 Ultra 550B A55B NVIDIA | 37.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 4, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 37.4 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.