MATH Level 5Benchmark scores and sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: math · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 GPT-5 OpenAI 98.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 29, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
98.1 100% confidence 100 percent, Full
2 GPT-5 OpenAI 97.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 20, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.9 100% confidence 100 percent, Full
3 GPT-5 Mini OpenAI 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 30, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.8 99% confidence 99 percent, High
4 o4-mini OpenAI 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.8 87% confidence 87 percent, High
5 o3 OpenAI 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.8 87% confidence 87 percent, High
6 Claude Sonnet 4.5 Anthropic 97.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 21, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.7 80% confidence 80 percent, High
7 Qwen3 Max Alibaba / Qwen 97.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 9, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.1 69% confidence 69 percent, Medium
8 GPT-5 Mini OpenAI 96.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 20, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
96.8 99% confidence 99 percent, High
9 DeepSeek-R1 DeepSeek 96.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 29, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
96.6 88% confidence 88 percent, High
10 Claude Haiku 4.5 Anthropic 96.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
96.4 64% confidence 64 percent, Medium
11 GPT-5 Nano OpenAI 95.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 20, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
95.2 85% confidence 85 percent, High
12 GPT-5 Nano OpenAI 94.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 20, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
94.9 85% confidence 85 percent, High
13 Claude Sonnet 3.7 Anthropic 91.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Mar 13, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.2 100% confidence 100 percent, Full
14 Claude Sonnet 3.7 Anthropic 90.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Mar 12, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
90.0 100% confidence 100 percent, Full
15 GPT-4.1 mini OpenAI 87.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.3 87% confidence 87 percent, High
16 Claude Haiku 4.5 Anthropic 86.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 16, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
86.9 64% confidence 64 percent, Medium
17 Claude Sonnet 3.7 Anthropic 86.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Feb 26, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
86.3 100% confidence 100 percent, Full
18 Claude Opus 4 Anthropic 85.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
85.0 100% confidence 100 percent, Full
19 Claude Sonnet 4 Anthropic 84.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
84.4 100% confidence 100 percent, Full
20 GPT-4.1 OpenAI 83.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
83.0 100% confidence 100 percent, Full
21 Mistral Medium 3 Mistral AI 81.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 7, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
81.6 85% confidence 85 percent, High
22 DeepSeek V3 0324 DeepSeek 75.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 1, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
75.5 74% confidence 74 percent, Medium
23 Gemma 3 27B IT Google 74.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Mar 13, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
74.0 74% confidence 74 percent, Medium
24 Llama 4 Maverick 17B Instruct Meta 73.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
73.0 78% confidence 78 percent, Medium
25 GPT-4.1 nano OpenAI 70.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
70.0 82% confidence 82 percent, High
26 Qwen3 235B-A22B Alibaba / Qwen 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 3, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
68.9 76% confidence 76 percent, Medium
27 Claude Sonnet 3.7 Anthropic 68.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 24, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
68.2 100% confidence 100 percent, Full
28 DeepSeek-V3 DeepSeek 64.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
64.9 93% confidence 93 percent, High
29 Llama 4 Scout 17B Instruct Meta 62.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
62.3 64% confidence 64 percent, Medium
30 Claude Sonnet 3.5 v2 Anthropic 56.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
56.9 100% confidence 100 percent, Full
31 Qwen Turbo Alibaba / Qwen 56.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 7, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
56.2 47% confidence 47 percent, Low
32 GPT-4o (2024-08-06) OpenAI 53.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
53.3 100% confidence 100 percent, Full
33 GPT-4o mini OpenAI 52.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
52.6 100% confidence 100 percent, Full
34 GPT-4o (2024-05-13) OpenAI 51.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
51.0 93% confidence 93 percent, High
35 Mistral Large 2.1 Mistral AI 50.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
50.3 100% confidence 100 percent, Full
36 GPT-4o (2024-11-20) OpenAI 49.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 5, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
49.8 61% confidence 61 percent, Medium
37 Mistral Small 3.1 24B Mistral AI 46.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Mar 18, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
46.8 64% confidence 64 percent, Medium
38 Claude Haiku 3.5 Anthropic 46.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Mar 12, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
46.4 100% confidence 100 percent, Full
39 Llama-3.3-70B-Instruct Meta 41.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
41.6 100% confidence 100 percent, Full
40 Llama-3.1-70B-Instruct Meta 36.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
36.7 93% confidence 93 percent, High
41 Llama-3.1-8B-Instruct Meta 22.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
22.9 93% confidence 93 percent, High
42 Claude Haiku 3 Anthropic 14.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
14.9 93% confidence 93 percent, High

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the math pillar of the SI Score.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed