Humanity's Last Exam (with tools)Benchmark scores and sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 Claude Opus 5.5 Anthropic 67.7%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; with tools; production safeguards with fallbackPublished Sep 22, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
67.7 100% confidence 100 percent, Full
2 Claude Fable 5.1 Anthropic 65%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools; production safeguards with fallbackPublished Sep 1, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
65.0 100% confidence 100 percent, Full
3 Claude Opus 5 Anthropic 64.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jul 24, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
64.7 100% confidence 100 percent, Full
4 Claude Fable 5 Anthropic 64.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
64.5 100% confidence 100 percent, Full
5 Claude Sonnet 5.5 Anthropic 64.5%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Sep 28, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
64.5 93% confidence 93 percent, High
6 DeepSeek V4.1 Flash DeepSeek 63.9%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] reasoning_effort=100; with tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
63.9 69% confidence 69 percent, Medium
7 GLM-5.3 Z.ai 62.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; with tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
62.5 93% confidence 93 percent, High
8 Muse Spark 1.1 Meta 62.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jul 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
62.1 61% confidence 61 percent, Medium
9 DeepSeek V4 Pro 0813 DeepSeek 60%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
60.0 69% confidence 69 percent, Medium
10 GPT-5.4 Pro OpenAI 58.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
58.7 69% confidence 69 percent, Medium
11 Claude Opus 4.8 Anthropic 57.9%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
57.9 100% confidence 100 percent, Full
12 Claude Haiku 5.5 Anthropic 57.4%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
57.4 68% confidence 68 percent, Medium
13 GPT-5.5 Pro OpenAI 57.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
57.2 53% confidence 53 percent, Medium
14 GPT-6 Astra OpenAI 57.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Sep 3, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
57.2 100% confidence 100 percent, Full
15 Qwen3.8 Max Alibaba / Qwen 56.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
56.2 85% confidence 85 percent, High
16 Qwen3.8 Max Preview Alibaba / Qwen 56.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] xhigh, with toolsPublished Aug 3, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
56.2 6% confidence 6 percent, Low
17 Claude Opus 4.7 Anthropic 54.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
54.7 100% confidence 100 percent, Full
18 GPT-5.5 OpenAI 52.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
52.2 100% confidence 100 percent, Full
19 GPT-5.4 OpenAI 52.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
52.1 100% confidence 100 percent, Full
20 DeepSeek V4 Pro DeepSeek 48.2%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; with tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
48.2 85% confidence 85 percent, High
21 Step 3.7 Flash StepFun 47.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished May 29, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
47.2 13% confidence 13 percent, Low
22 Claude Sonnet 4.6 Anthropic 46.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 30, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
46.8 100% confidence 100 percent, Full
23 DeepSeek V4 Flash DeepSeek 45.1%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort; with tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
45.1 53% confidence 53 percent, Medium
24 GPT-5.4 mini OpenAI 41.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Mar 17, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
41.5 100% confidence 100 percent, Full
25 GPT-5.4 nano OpenAI 37.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Mar 17, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
37.7 100% confidence 100 percent, Full
26 Nemotron 3 Ultra 550B A55B NVIDIA 37.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 4, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
37.4 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed