FrontierMath Tiers 1–3 (v2)Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | GPT-6 Astra OpenAI | 93.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.7 | 100% confidence 100 percent, Full |
| 2 | GPT-6.1 Sol OpenAI | 93.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.7 | 100% confidence 100 percent, Full |
| 3 | Claude Opus 5.5 Anthropic | 91.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.2 | 100% confidence 100 percent, Full |
| 4 | Claude Fable 5.1 Anthropic | 90.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 1, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.2 | 100% confidence 100 percent, Full |
| 5 | GPT-6 Sol OpenAI | 89.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 89.8 | 100% confidence 100 percent, Full |
| 6 | GPT-5.6 Sol OpenAI | 89.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 89.1 | 100% confidence 100 percent, Full |
| 7 | Claude Sonnet 5.5 Anthropic | 88.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.8 | 93% confidence 93 percent, High |
| 8 | GPT-5.5 Pro OpenAI | 87.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.7 | 53% confidence 53 percent, Medium |
| 9 | Claude Fable 5 Anthropic | 87.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.0 | 100% confidence 100 percent, Full |
| 10 | GPT-5.6 Terra OpenAI | 86.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.0 | 100% confidence 100 percent, Full |
| 11 | Claude Opus 5 Anthropic | 85.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.6 | 100% confidence 100 percent, Full |
| 12 | GPT-5.5 OpenAI | 85.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.3 | 100% confidence 100 percent, Full |
| 13 | GPT-5.4 Pro OpenAI | 82.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.5 | 69% confidence 69 percent, Medium |
| 14 | GPT-5.6 Luna OpenAI | 82.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.1 | 100% confidence 100 percent, Full |
| 15 | Claude Opus 4.8 Anthropic | 80%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.0 | 100% confidence 100 percent, Full |
| 16 | GPT-6 Luna OpenAI | 78.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.9 | 100% confidence 100 percent, Full |
| 17 | GPT-5.4 OpenAI | 78.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.6 | 100% confidence 100 percent, Full |
| 18 | Qwen3.8 Max Alibaba / Qwen | 74.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 74.7 | 85% confidence 85 percent, High |
| 19 | Muse Spark 1.3 Meta | 74.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 16, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 74.4 | 93% confidence 93 percent, High |
| 20 | Muse Spark 1.3 Meta | 74.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 18, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 74.0 | 93% confidence 93 percent, High |
| 21 | GPT-5.2 Pro OpenAI | 74%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 74.0 | 48% confidence 48 percent, Low |
| 22 | Kimi K3 Moonshot AI | 72.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 72.2 | 100% confidence 100 percent, Full |
| 23 | Gemini 3.7 Flash Google | 71.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.6 | 100% confidence 100 percent, Full |
| 24 | Claude Opus 4.7 Anthropic | 70.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 70.2 | 100% confidence 100 percent, Full |
| 25 | GLM-5.3 Z.ai | 68.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 25, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.8 | 93% confidence 93 percent, High |
| 26 | Gemini 3.8 Flash Google | 68.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.4 | 100% confidence 100 percent, Full |
| 27 | GPT-5.2 OpenAI | 67.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 67.4 | 100% confidence 100 percent, Full |
| 28 | Claude Opus 4.6 Anthropic | 66.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.0 | 100% confidence 100 percent, Full |
| 29 | Grok 4.6 xAI | 66.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.0 | 100% confidence 100 percent, Full |
| 30 | Qwen3.8 Max 0902 Alibaba / Qwen | 65.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 65.6 | 40% confidence 40 percent, Low |
| 31 | Claude Sonnet 5 Anthropic | 65.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 65.6 | 93% confidence 93 percent, High |
| 32 | Qwen3.7 Max Alibaba / Qwen | 64.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 64.6 | 53% confidence 53 percent, Medium |
| 33 | DeepSeek V4 Pro 0813 DeepSeek | 64.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 19, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 64.6 | 69% confidence 69 percent, Medium |
| 34 | Gemini 3.5 Flash Google | 62.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 62.8 | 100% confidence 100 percent, Full |
| 35 | Gemini 3.1 Pro Preview Google | 59.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 59.6 | 100% confidence 100 percent, Full |
| 36 | GLM-5.2 Z.ai | 59.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 19, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 59.2 | 100% confidence 100 percent, Full |
| 37 | Gemini 3.6 Flash Google | 58.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 58.9 | 100% confidence 100 percent, Full |
| 38 | DeepSeek V4 Flash 0731 DeepSeek | 57.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 57.5 | 69% confidence 69 percent, Medium |
| 39 | Kimi K2.6 Moonshot AI | 57.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 57.2 | 85% confidence 85 percent, High |
| 40 | Grok 4.5 xAI | 57.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 57.2 | 100% confidence 100 percent, Full |
| 41 | GPT-5 Pro OpenAI | 55.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 55.8 | 64% confidence 64 percent, Medium |
| 42 | GLM-5.3-Flash Z.ai | 55.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 55.8 | 100% confidence 100 percent, Full |
| 43 | GPT-5 OpenAI | 55.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 55.4 | 100% confidence 100 percent, Full |
| 44 | GLM-5.2 Z.ai | 54.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 54.7 | 100% confidence 100 percent, Full |
| 45 | Kimi K2.7 Code Moonshot AI | 54.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 54.0 | 48% confidence 48 percent, Low |
| 46 | Grok 4.7 xAI | 53.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 53.0 | 100% confidence 100 percent, Full |
| 47 | Gemini 3 Flash Preview Google | 51.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 51.2 | 93% confidence 93 percent, High |
| 48 | GPT-5.4 mini OpenAI | 51.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 51.2 | 100% confidence 100 percent, Full |
| 49 | GPT-5 Mini OpenAI | 46.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.7 | 99% confidence 99 percent, High |
| 50 | Inkling Small Thinking Machines Lab | 46.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.3 | 100% confidence 100 percent, Full |
| 51 | DeepSeek V4 Pro DeepSeek | 45.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 45.3 | 85% confidence 85 percent, High |
| 52 | GPT-5.4 nano OpenAI | 44.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 44.9 | 100% confidence 100 percent, Full |
| 53 | Grok 4.20 (Reasoning) xAI | 44.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 44.9 | 80% confidence 80 percent, High |
| 54 | Grok 4.3 xAI | 42.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 42.8 | 80% confidence 80 percent, High |
| 55 | GLM-5.2 Z.ai | 42.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 42.5 | 100% confidence 100 percent, Full |
| 56 | GPT-5.6 Luna OpenAI | 41.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 41.4 | 100% confidence 100 percent, Full |
| 57 | GPT-5.6 Luna OpenAI | 39.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 39.6 | 100% confidence 100 percent, Full |
| 58 | Qwen3.6 Plus Alibaba / Qwen | 38.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 38.2 | 80% confidence 80 percent, High |
| 59 | GPT-5 OpenAI | 37.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 37.2 | 100% confidence 100 percent, Full |
| 60 | GLM-5.1 Z.ai | 36.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 36.8 | 64% confidence 64 percent, Medium |
| 61 | o4-mini OpenAI | 36.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 36.1 | 87% confidence 87 percent, High |
| 62 | Qwen3.6 27B Alibaba / Qwen | 35.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 35.1 | 53% confidence 53 percent, Medium |
| 63 | Qwen3.7 Plus Alibaba / Qwen | 34.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 34.4 | 64% confidence 64 percent, Medium |
| 64 | Claude Opus 4.5 Anthropic | 34.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 34.4 | 100% confidence 100 percent, Full |
| 65 | Qwen3.6 27B Alibaba / Qwen | 34.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 34.0 | 53% confidence 53 percent, Medium |
| 66 | o3 OpenAI | 33.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 33.3 | 87% confidence 87 percent, High |
| 67 | Inkling Thinking Machines Lab | 33.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 33.3 | 100% confidence 100 percent, Full |
| 68 | Qwen3.6 Plus Alibaba / Qwen | 32.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 32.3 | 80% confidence 80 percent, High |
| 69 | Qwen3.5 397B-A17B Alibaba / Qwen | 31.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 31.2 | 69% confidence 69 percent, Medium |
| 70 | o3 OpenAI | 29.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 29.8 | 87% confidence 87 percent, High |
| 71 | Qwen3.5 397B-A17B Alibaba / Qwen | 29.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 29.5 | 69% confidence 69 percent, Medium |
| 72 | o4-mini OpenAI | 28.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 28.8 | 87% confidence 87 percent, High |
| 73 | Gemini 3.1 Flash Lite Google | 27.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 27.7 | 32% confidence 32 percent, Low |
| 74 | GPT-5.5 Instant OpenAI | 26.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 26.3 | 64% confidence 64 percent, Medium |
| 75 | Gemini 3.5 Flash Lite Google | 26.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 26.0 | 100% confidence 100 percent, Full |
| 76 | GLM-5.1 Z.ai | 24.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 24.9 | 64% confidence 64 percent, Medium |
| 77 | Gemini 2.5 Pro Google | 24.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 24.6 | 90% confidence 90 percent, High |
| 78 | GPT-5.4 mini OpenAI | 24.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 24.6 | 100% confidence 100 percent, Full |
| 79 | Claude Sonnet 4.5 Anthropic | 23.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 23.9 | 80% confidence 80 percent, High |
| 80 | Qwen3.6 Flash Alibaba / Qwen | 22.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 22.5 | 32% confidence 32 percent, Low |
| 81 | Gemini 3.1 Flash Lite Google | 22.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 22.5 | 32% confidence 32 percent, Low |
| 82 | Gemini 3.1 Flash Lite Google | 21.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 21.4 | 32% confidence 32 percent, Low |
| 83 | Qwen3.6 35B-A3B Alibaba / Qwen | 20.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 20.4 | 37% confidence 37 percent, Low |
| 84 | GPT-5.4 nano OpenAI | 20.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 20.4 | 100% confidence 100 percent, Full |
| 85 | GPT-5 Nano OpenAI | 20%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 20.0 | 85% confidence 85 percent, High |
| 86 | Qwen3.7 Flash Alibaba / Qwen | 19.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 19.3 | 32% confidence 32 percent, Low |
| 87 | Qwen3.7 Flash Alibaba / Qwen | 19.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 19.3 | 32% confidence 32 percent, Low |
| 88 | o3 OpenAI | 19.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 19.3 | 87% confidence 87 percent, High |
| 89 | Qwen3 Max Alibaba / Qwen | 18.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 18.9 | 69% confidence 69 percent, Medium |
| 90 | Qwen3.5 Flash Alibaba / Qwen | 18.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 18.2 | 64% confidence 64 percent, Medium |
| 91 | GPT-5 OpenAI | 18.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 18.2 | 100% confidence 100 percent, Full |
| 92 | GPT-5 Mini OpenAI | 18.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 18.2 | 99% confidence 99 percent, High |
| 93 | Qwen3.6 35B-A3B Alibaba / Qwen | 17.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 17.5 | 37% confidence 37 percent, Low |
| 94 | Qwen3.6 Flash Alibaba / Qwen | 17.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 17.2 | 32% confidence 32 percent, Low |
| 95 | GPT-5.4 mini OpenAI | 17.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 17.2 | 100% confidence 100 percent, Full |
| 96 | o4-mini OpenAI | 16.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 16.1 | 87% confidence 87 percent, High |
| 97 | Claude Opus 4.1 Anthropic | 12.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 12.6 | 80% confidence 80 percent, High |
| 98 | Qwen3.5 Flash Alibaba / Qwen | 9.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 9.5 | 64% confidence 64 percent, Medium |
| 99 | GPT-4.1 mini OpenAI | 6.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 6.7 | 87% confidence 87 percent, High |
| 100 | GPT-4.1 OpenAI | 6.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 6.0 | 100% confidence 100 percent, Full |
| 101 | GPT-5 Mini OpenAI | 6.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 6.0 | 99% confidence 99 percent, High |
| 102 | GPT-5 Nano OpenAI | 6.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 6.0 | 85% confidence 85 percent, High |
| 103 | GPT-5.4 nano OpenAI | 4.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 4.6 | 100% confidence 100 percent, Full |
| 104 | GPT-5 Nano OpenAI | 1.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 1.8 | 85% confidence 85 percent, High |
| 105 | GPT-4o mini OpenAI | 0.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 0.7 | 100% confidence 100 percent, Full |
| 106 | GPT-4o (2024-08-06) OpenAI | 0.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 0.4 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the math pillar of the SI Score.