Humanity's Last ExamBenchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 Anthropic | 60.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools; production safeguards with fallbackPublished Sep 1, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 60.9 | 100% confidence 100 percent, Full |
| 2 | Claude Fable 5 Anthropic | 59%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 59.0 | 100% confidence 100 percent, Full |
| 3 | Claude Opus 5 Anthropic | 56.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jul 24, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 56.3 | 100% confidence 100 percent, Full |
| 4 | Fugu Ultra Sakana AI | 50%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant]
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 50.0 | 100% confidence 100 percent, Full |
| 5 | Claude Opus 4.8 Anthropic | 49.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 9, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 49.8 | 100% confidence 100 percent, Full |
| 6 | Fugu Sakana AI | 47.2%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant]
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 47.2 | 100% confidence 100 percent, Full |
| 7 | Claude Opus 4.7 Anthropic | 46.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 46.9 | 100% confidence 100 percent, Full |
| 8 | Claude Haiku 5.5 Anthropic | 45.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 45.9 | 68% confidence 68 percent, Medium |
| 9 | Qwen3.8 Max Alibaba / Qwen | 43.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 43.6 | 85% confidence 85 percent, High |
| 10 | Qwen3.8 Max Preview Alibaba / Qwen | 43.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh, no toolsPublished Aug 3, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 43.6 | 6% confidence 6 percent, Low |
| 11 | GPT-5.5 Pro OpenAI | 43.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 43.1 | 53% confidence 53 percent, Medium |
| 12 | DeepSeek V4 Pro 0813 DeepSeek | 42.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 42.7 | 69% confidence 69 percent, Medium |
| 13 | GPT-5.4 Pro OpenAI | 42.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 42.7 | 69% confidence 69 percent, Medium |
| 14 | Qwen3.7 Max Alibaba / Qwen | 41.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published May 19, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 41.4 | 53% confidence 53 percent, Medium |
| 15 | GPT-5.5 OpenAI | 41.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 41.4 | 100% confidence 100 percent, Full |
| 16 | GPT-5.4 OpenAI | 39.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 39.8 | 100% confidence 100 percent, Full |
| 17 | DeepSeek V4 Pro DeepSeek | 37.7%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; without tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 37.7 | 85% confidence 85 percent, High |
| 18 | Qwen3.8 Flash Next Alibaba / Qwen | 35.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools; GPT-4o judge
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 35.9 | 21% confidence 21 percent, Low |
| 19 | DeepSeek V4 Flash DeepSeek | 34.8%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; without tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 34.8 | 53% confidence 53 percent, Medium |
| 20 | Claude Sonnet 4.6 Anthropic | 34.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 30, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 34.6 | 100% confidence 100 percent, Full |
| 21 | Qwen3.8 27B Alibaba / Qwen | 30.8%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools; GPT-4o judge
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 30.8 | 69% confidence 69 percent, Medium |
| 22 | GPT-5.4 mini OpenAI | 28.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] without toolsPublished Mar 17, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 28.2 | 100% confidence 100 percent, Full |
| 23 | Nemotron 3 Ultra 550B A55B NVIDIA | 26.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 4, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 26.7 | 100% confidence 100 percent, Full |
| 24 | GPT-5.4 nano OpenAI | 24.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] without toolsPublished Mar 17, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 24.3 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.