FrontierMath Tier 4 (v2)Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | GPT-6.1 Sol OpenAI | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 100% confidence 100 percent, Full |
| 2 | GPT-6 Astra OpenAI | 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.6 | 100% confidence 100 percent, Full |
| 3 | GPT-6 Astra OpenAI | 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.6 | 100% confidence 100 percent, Full |
| 4 | GPT-6 Astra OpenAI | 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.6 | 100% confidence 100 percent, Full |
| 5 | GPT-6 Astra OpenAI | 97.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.6 | 100% confidence 100 percent, Full |
| 6 | Claude Opus 5.5 Anthropic | 95%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.0 | 100% confidence 100 percent, Full |
| 7 | Claude Fable 5 Anthropic | 90.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.2 | 100% confidence 100 percent, Full |
| 8 | GPT-6 Sol OpenAI | 90%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 9 | Claude Fable 5.1 Anthropic | 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 1, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.8 | 100% confidence 100 percent, Full |
| 10 | GPT-6 Astra OpenAI | 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.8 | 100% confidence 100 percent, Full |
| 11 | GPT-5.6 Sol OpenAI | 82.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.9 | 100% confidence 100 percent, Full |
| 12 | GPT-6 Astra OpenAI | 82.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.9 | 100% confidence 100 percent, Full |
| 13 | Claude Sonnet 5.5 Anthropic | 80.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.5 | 93% confidence 93 percent, High |
| 14 | GPT-5.6 Sol OpenAI | 80.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] promaxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.5 | 100% confidence 100 percent, Full |
| 15 | GPT-5.5 Pro OpenAI | 78.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.0 | 53% confidence 53 percent, Medium |
| 16 | Claude Opus 5 Anthropic | 73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.2 | 100% confidence 100 percent, Full |
| 17 | GPT-5.5 OpenAI | 72.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 72.5 | 100% confidence 100 percent, Full |
| 18 | GPT-5.6 Terra OpenAI | 70.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 70.7 | 100% confidence 100 percent, Full |
| 19 | GPT-5.6 Luna OpenAI | 61.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 61.0 | 100% confidence 100 percent, Full |
| 20 | GPT-5.4 Pro OpenAI | 58.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 58.5 | 69% confidence 69 percent, Medium |
| 21 | Claude Opus 4.8 Anthropic | 56.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 56.1 | 100% confidence 100 percent, Full |
| 22 | GPT-6 Luna OpenAI | 56.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 56.1 | 100% confidence 100 percent, Full |
| 23 | GPT-5.4 OpenAI | 49%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 49.0 | 100% confidence 100 percent, Full |
| 24 | Qwen3.8 Max Alibaba / Qwen | 46.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.3 | 85% confidence 85 percent, High |
| 25 | Muse Spark 1.3 Meta | 46.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 18, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.3 | 93% confidence 93 percent, High |
| 26 | GPT-5.2 Pro OpenAI | 46%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.0 | 48% confidence 48 percent, Low |
| 27 | Muse Spark 1.3 Meta | 41.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 16, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 41.5 | 93% confidence 93 percent, High |
| 28 | Kimi K3 Moonshot AI | 39.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 39.0 | 100% confidence 100 percent, Full |
| 29 | Gemini 3.7 Flash Google | 36.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 36.6 | 100% confidence 100 percent, Full |
| 30 | Qwen3.7 Max Alibaba / Qwen | 34.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 34.1 | 53% confidence 53 percent, Medium |
| 31 | Qwen3.8 Max 0902 Alibaba / Qwen | 34.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 34.1 | 40% confidence 40 percent, Low |
| 32 | Claude Opus 4.7 Anthropic | 31.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 31.7 | 100% confidence 100 percent, Full |
| 33 | Grok 4.6 xAI | 31.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 31.7 | 100% confidence 100 percent, Full |
| 34 | GPT-5.2 OpenAI | 31.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 31.7 | 100% confidence 100 percent, Full |
| 35 | Claude Sonnet 5 Anthropic | 29.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 29.3 | 93% confidence 93 percent, High |
| 36 | GLM-5.2 Z.ai | 29.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 19, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 29.3 | 100% confidence 100 percent, Full |
| 37 | GLM-5.3 Z.ai | 29.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 25, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 29.3 | 93% confidence 93 percent, High |
| 38 | Claude Opus 4.6 Anthropic | 26.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 26.8 | 100% confidence 100 percent, Full |
| 39 | DeepSeek V4 Pro 0813 DeepSeek | 26.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 19, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 26.8 | 69% confidence 69 percent, Medium |
| 40 | Gemini 3.1 Pro Preview Google | 26.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 26.8 | 100% confidence 100 percent, Full |
| 41 | Gemini 3.5 Flash Google | 26.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 26.8 | 100% confidence 100 percent, Full |
| 42 | Kimi K2.6 Moonshot AI | 25.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 25.6 | 85% confidence 85 percent, High |
| 43 | DeepSeek V4 Flash 0731 DeepSeek | 24.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 24.4 | 69% confidence 69 percent, Medium |
| 44 | Grok 4.5 xAI | 24.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 24.4 | 100% confidence 100 percent, Full |
| 45 | Gemini 3.6 Flash Google | 22.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 22.0 | 100% confidence 100 percent, Full |
| 46 | Gemini 3.8 Flash Google | 22.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 22.0 | 100% confidence 100 percent, Full |
| 47 | GPT-5 OpenAI | 22.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 22.0 | 100% confidence 100 percent, Full |
| 48 | GPT-5 Pro OpenAI | 19.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 19.5 | 64% confidence 64 percent, Medium |
| 49 | Gemini 3 Flash Preview Google | 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 17.1 | 93% confidence 93 percent, High |
| 50 | Inkling Small Thinking Machines Lab | 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 17.1 | 100% confidence 100 percent, Full |
| 51 | Grok 4.20 (Reasoning) xAI | 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 17.1 | 80% confidence 80 percent, High |
| 52 | Grok 4.7 xAI | 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 17.1 | 100% confidence 100 percent, Full |
| 53 | GLM-5.3-Flash Z.ai | 17.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 17.1 | 100% confidence 100 percent, Full |
| 54 | Grok 4.3 xAI | 14.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 14.6 | 80% confidence 80 percent, High |
| 55 | Kimi K2.7 Code Moonshot AI | 12.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 12.2 | 48% confidence 48 percent, Low |
| 56 | GPT-5 Mini OpenAI | 12.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 12.2 | 99% confidence 99 percent, High |
| 57 | GPT-5.4 nano OpenAI | 12.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 12.2 | 100% confidence 100 percent, Full |
| 58 | GPT-5.4 mini OpenAI | 9.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 9.8 | 100% confidence 100 percent, Full |
| 59 | Claude Opus 4.5 Anthropic | 4.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 4.9 | 100% confidence 100 percent, Full |
| 60 | o4-mini OpenAI | 4.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 4.9 | 87% confidence 87 percent, High |
| 61 | Inkling Thinking Machines Lab | 4.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 4.9 | 100% confidence 100 percent, Full |
| 62 | Claude Opus 4.1 Anthropic | 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 2.4 | 80% confidence 80 percent, High |
| 63 | Claude Sonnet 4.5 Anthropic | 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 2.4 | 80% confidence 80 percent, High |
| 64 | DeepSeek V4 Pro DeepSeek | 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 2.4 | 85% confidence 85 percent, High |
| 65 | GPT-5 Nano OpenAI | 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 2.4 | 85% confidence 85 percent, High |
| 66 | GPT-5.5 Instant OpenAI | 2.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 2.4 | 64% confidence 64 percent, Medium |
| 67 | Gemini 2.5 Pro Google | 0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jun 11, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 0.0 | 90% confidence 90 percent, High |
| 68 | Gemini 3.5 Flash Lite Google | 0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the math pillar of the SI Score.