Humanity's Last Exam (Scale AI)Benchmark scores and sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 GPT-6 Astra OpenAI 54.8%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Sep 9, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
54.8 100% confidence 100 percent, Full
2 Claude Fable 5.1 Anthropic 46.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Sep 3, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
46.5 100% confidence 100 percent, Full
3 Gemini 3.1 Pro Preview Google 46.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
46.4 100% confidence 100 percent, Full
4 Gemini 3.8 Flash Google 44.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Sep 9, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
44.5 100% confidence 100 percent, Full
5 GPT-5.4 Pro OpenAI 44.3%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Mar 23, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
44.3 69% confidence 69 percent, Medium
6 Gemini 3 Pro Preview Google 37.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Nov 19, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
37.5 68% confidence 68 percent, Medium
7 GPT-5.4 OpenAI 36.2%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Mar 10, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
36.2 100% confidence 100 percent, Full
8 Claude Opus 4.7 Anthropic 36.2%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 22, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
36.2 100% confidence 100 percent, Full
9 GPT-5 Pro OpenAI 31.6%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Nov 6, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
31.6 64% confidence 64 percent, Medium
10 Claude Haiku 5.5 Anthropic 30.1%Humanity’s Last ExamPublished steward score [variant] claude-haiku-5-5Published Oct 8, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
30.1 68% confidence 68 percent, Medium
11 GPT-5.2 OpenAI 27.8%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Dec 15, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
27.8 100% confidence 100 percent, Full
12 GPT-5 OpenAI 25.3%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. Sampled at reasoning_effort: 'high'. [variant] Published Aug 7, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
25.3 100% confidence 100 percent, Full
13 Kimi K2.5 Moonshot AI 24.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Feb 13, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
24.4 100% confidence 100 percent, Full
14 GPT-5 Mini OpenAI 19.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Aug 22, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
19.4 99% confidence 99 percent, High
15 Claude Opus 4.6 Anthropic 19%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Feb 17, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
19.0 100% confidence 100 percent, Full
16 Claude Opus 4.5 Anthropic 14.2%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Nov 26, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
14.2 100% confidence 100 percent, Full
17 Claude Opus 4 Anthropic 10.7%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published May 24, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
10.7 100% confidence 100 percent, Full
18 Gemini 3.1 Flash Lite Preview Google 8.6%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Mar 23, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
8.6 48% confidence 48 percent, Low
19 o1-pro OpenAI 8.1%Humanity’s Last Exam9% (216 prompts) failed due to a post-training bug and were counted as failures. OpenAI has been informed and is working on a fix. --- Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
8.1 16% confidence 16 percent, Low
20 Claude Sonnet 3.7 Anthropic 8.0%Humanity’s Last ExamThinking budget: 16,000 tokens. Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
8.0 100% confidence 100 percent, Full
21 Claude Opus 4.1 Anthropic 7.9%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Aug 8, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
7.9 80% confidence 80 percent, High
22 Claude Sonnet 4.5 Anthropic 7.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Oct 2, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
7.5 80% confidence 80 percent, High
23 Llama 4 Maverick 17B Instruct Meta 5.7%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.7 78% confidence 78 percent, Medium
24 Claude Sonnet 4 Anthropic 5.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published May 23, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.5 100% confidence 100 percent, Full
25 GPT-4.1 OpenAI 5.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 14, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.4 100% confidence 100 percent, Full
26 Mistral Medium 3 Mistral AI 4.5%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published May 13, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.5 85% confidence 85 percent, High
27 Nova Pro Amazon 4.4%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.4 100% confidence 100 percent, Full
28 Nova Lite Amazon 3.6%Humanity’s Last ExamPotential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
3.6 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed