OTIS Mock AIME 2024–2025Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Qwen3.8 Max 0902 Alibaba / Qwen | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 40% confidence 40 percent, Low |
| 2 | Claude Fable 5 Anthropic | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 100% confidence 100 percent, Full |
| 3 | Claude Fable 5.1 Anthropic | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 1, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 100% confidence 100 percent, Full |
| 4 | Claude Opus 5.5 Anthropic | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 100% confidence 100 percent, Full |
| 5 | Claude Sonnet 5.5 Anthropic | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 93% confidence 93 percent, High |
| 6 | GPT-5.6 Sol OpenAI | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 100% confidence 100 percent, Full |
| 7 | GPT-6 Astra OpenAI | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 100% confidence 100 percent, Full |
| 8 | GPT-6 Sol OpenAI | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 100% confidence 100 percent, Full |
| 9 | GPT-6.1 Sol OpenAI | 100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 100.0 | 100% confidence 100 percent, Full |
| 10 | Claude Fable 5 Anthropic | 99.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 99.7 | 100% confidence 100 percent, Full |
| 11 | GPT-5.6 Terra OpenAI | 99.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 99.7 | 100% confidence 100 percent, Full |
| 12 | Qwen3.8 Max Alibaba / Qwen | 99.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 99.4 | 85% confidence 85 percent, High |
| 13 | Muse Spark 1.3 Meta | 99.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 99.2 | 93% confidence 93 percent, High |
| 14 | Grok 4.6 xAI | 99.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 99.2 | 100% confidence 100 percent, Full |
| 15 | Claude Opus 5 Anthropic | 98.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 98.9 | 100% confidence 100 percent, Full |
| 16 | Gemini 3.8 Flash Google | 98.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Sep 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 98.9 | 100% confidence 100 percent, Full |
| 17 | GPT-6 Luna OpenAI | 98.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 98.9 | 100% confidence 100 percent, Full |
| 18 | DeepSeek V4 Pro 0813 DeepSeek | 98.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 18, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 98.6 | 69% confidence 69 percent, Medium |
| 19 | Claude Opus 4.8 Anthropic | 98.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 98.3 | 100% confidence 100 percent, Full |
| 20 | GPT-5.6 Luna OpenAI | 98.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 9, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 98.3 | 100% confidence 100 percent, Full |
| 21 | Grok 4.7 xAI | 98.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Sep 22, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 98.1 | 100% confidence 100 percent, Full |
| 22 | Claude Opus 4.7 Anthropic | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Apr 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 23 | Claude Fable 5 Anthropic | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 24 | Claude Opus 4.8 Anthropic | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 25 | Claude Opus 5 Anthropic | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 26 | GPT-5.4 OpenAI | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 27 | Grok 4.5 xAI | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 8, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 28 | Grok 4.6 xAI | 97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 29 | Gemini 3.7 Flash Google | 97.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.2 | 100% confidence 100 percent, Full |
| 30 | Kimi K3 Moonshot AI | 97.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 16, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 97.2 | 100% confidence 100 percent, Full |
| 31 | DeepSeek V4 Pro DeepSeek | 96.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 96.7 | 85% confidence 85 percent, High |
| 32 | Kimi K2.6 Moonshot AI | 96.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 96.1 | 85% confidence 85 percent, High |
| 33 | GPT-5.2 OpenAI | 96.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Dec 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 96.1 | 100% confidence 100 percent, Full |
| 34 | GPT-5.2 OpenAI | 96.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Dec 13, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 96.1 | 100% confidence 100 percent, Full |
| 35 | Gemini 3.1 Pro Preview Google | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 20, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 100% confidence 100 percent, Full |
| 36 | Qwen3.7 Max Alibaba / Qwen | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 53% confidence 53 percent, Medium |
| 37 | DeepSeek V4 Pro DeepSeek | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 85% confidence 85 percent, High |
| 38 | Gemini 3 Flash Preview Google | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 93% confidence 93 percent, High |
| 39 | Gemini 3.1 Pro Preview Google | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 100% confidence 100 percent, Full |
| 40 | Gemini 3.5 Flash Google | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished May 25, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 100% confidence 100 percent, Full |
| 41 | Kimi K2.7 Code Moonshot AI | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 48% confidence 48 percent, Low |
| 42 | GPT-5.4 OpenAI | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 100% confidence 100 percent, Full |
| 43 | GPT-5.6 Sol OpenAI | 95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.6 | 100% confidence 100 percent, Full |
| 44 | GPT-5.4 OpenAI | 95.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Mar 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 95.3 | 100% confidence 100 percent, Full |
| 45 | Claude Sonnet 5 Anthropic | 94.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jul 1, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.7 | 93% confidence 93 percent, High |
| 46 | Claude Opus 4.6 Anthropic | 94.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Feb 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.4 | 100% confidence 100 percent, Full |
| 47 | DeepSeek V4 Flash 0731 DeepSeek | 94.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.4 | 69% confidence 69 percent, Medium |
| 48 | Gemini 3.6 Flash Google | 94.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 94.2 | 100% confidence 100 percent, Full |
| 49 | GPT-5.2 OpenAI | 93.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Dec 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.9 | 100% confidence 100 percent, Full |
| 50 | GLM-5.3-Flash Z.ai | 93.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 26, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.9 | 100% confidence 100 percent, Full |
| 51 | Qwen3.6 Plus Alibaba / Qwen | 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.3 | 80% confidence 80 percent, High |
| 52 | Qwen3.7 Plus Alibaba / Qwen | 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.3 | 64% confidence 64 percent, Medium |
| 53 | Claude Opus 5 Anthropic | 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.3 | 100% confidence 100 percent, Full |
| 54 | Kimi K3 Moonshot AI | 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.3 | 100% confidence 100 percent, Full |
| 55 | Grok 4.3 xAI | 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jun 17, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.3 | 80% confidence 80 percent, High |
| 56 | GLM-5.1 Z.ai | 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.3 | 64% confidence 64 percent, Medium |
| 57 | Claude Opus 4.6 Anthropic | 93.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Feb 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 93.1 | 100% confidence 100 percent, Full |
| 58 | Gemini 3 Flash Preview Google | 92.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Dec 17, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 92.8 | 93% confidence 93 percent, High |
| 59 | Grok 4.20 (Reasoning) xAI | 92.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 92.2 | 80% confidence 80 percent, High |
| 60 | Gemini 3 Pro Preview Google | 91.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 19, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.4 | 68% confidence 68 percent, Medium |
| 61 | GPT-5 OpenAI | 91.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 29, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.4 | 100% confidence 100 percent, Full |
| 62 | Qwen3.6 27B Alibaba / Qwen | 91.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.1 | 53% confidence 53 percent, Medium |
| 63 | Qwen3.6 Max Preview Alibaba / Qwen | 91.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.1 | 64% confidence 64 percent, Medium |
| 64 | Claude Opus 4.6 Anthropic | 91.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.1 | 100% confidence 100 percent, Full |
| 65 | GLM-5.3 Z.ai | 91.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 24, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 91.1 | 93% confidence 93 percent, High |
| 66 | Inkling Small Thinking Machines Lab | 90%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 67 | Qwen3.5 397B-A17B Alibaba / Qwen | 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.9 | 69% confidence 69 percent, Medium |
| 68 | Gemini 3.5 Flash Google | 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.9 | 100% confidence 100 percent, Full |
| 69 | GPT-5.4 mini OpenAI | 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.9 | 100% confidence 100 percent, Full |
| 70 | GPT-5.6 Terra OpenAI | 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.9 | 100% confidence 100 percent, Full |
| 71 | GPT OSS 120B OpenAI | 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Dec 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.9 | 79% confidence 79 percent, Medium |
| 72 | Inkling Thinking Machines Lab | 88.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 5, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.9 | 100% confidence 100 percent, Full |
| 73 | GPT-5.1 OpenAI | 88.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Nov 13, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 88.6 | 85% confidence 85 percent, High |
| 74 | DeepSeek Reasoner DeepSeek | 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Dec 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.8 | 32% confidence 32 percent, Low |
| 75 | GPT-5.4 nano OpenAI | 87.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 14, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.8 | 100% confidence 100 percent, Full |
| 76 | GPT-5 OpenAI | 87.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.2 | 100% confidence 100 percent, Full |
| 77 | GPT-5.4 mini OpenAI | 87.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 87.2 | 100% confidence 100 percent, Full |
| 78 | Qwen3.5 Plus Alibaba / Qwen | 86.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.7 | 32% confidence 32 percent, Low |
| 79 | Qwen3.6 35B-A3B Alibaba / Qwen | 86.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.7 | 37% confidence 37 percent, Low |
| 80 | Qwen3.7 Flash Alibaba / Qwen | 86.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.7 | 32% confidence 32 percent, Low |
| 81 | Claude Opus 4.7 Anthropic | 86.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.7 | 100% confidence 100 percent, Full |
| 82 | Nemotron 3 Ultra 550B A55B NVIDIA | 86.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.7 | 100% confidence 100 percent, Full |
| 83 | GPT-5 Mini OpenAI | 86.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 30, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.7 | 99% confidence 99 percent, High |
| 84 | GLM-5.2 Z.ai | 86.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 25, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.4 | 100% confidence 100 percent, Full |
| 85 | Claude Opus 4.5 Anthropic | 86.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Nov 24, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 86.1 | 100% confidence 100 percent, Full |
| 86 | Claude Sonnet 4.6 Anthropic | 85.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Feb 20, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.8 | 100% confidence 100 percent, Full |
| 87 | GPT-5.1 OpenAI | 85.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Nov 17, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 85.6 | 85% confidence 85 percent, High |
| 88 | Qwen3.5 Flash Alibaba / Qwen | 84.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.4 | 64% confidence 64 percent, Medium |
| 89 | Qwen3.6 Flash Alibaba / Qwen | 84.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.4 | 32% confidence 32 percent, Low |
| 90 | Claude Opus 4.8 Anthropic | 84.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.4 | 100% confidence 100 percent, Full |
| 91 | GPT-5.4 OpenAI | 84.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.4 | 100% confidence 100 percent, Full |
| 92 | GPT-5.5 OpenAI | 84.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.4 | 100% confidence 100 percent, Full |
| 93 | o3 OpenAI | 84.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.4 | 87% confidence 87 percent, High |
| 94 | Gemini 2.5 Pro Google | 84.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 84.2 | 90% confidence 90 percent, High |
| 95 | o3 OpenAI | 83.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.9 | 87% confidence 87 percent, High |
| 96 | GLM-4.7 Z.ai | 83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 29, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.3 | 69% confidence 69 percent, Medium |
| 97 | Kimi K2 Thinking Turbo Moonshot AI | 83.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 13, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 83.1 | 64% confidence 64 percent, Medium |
| 98 | Qwen3.5 397B-A17B Alibaba / Qwen | 82.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.2 | 69% confidence 69 percent, Medium |
| 99 | Claude Sonnet 4.6 Anthropic | 82.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.2 | 100% confidence 100 percent, Full |
| 100 | Gemini 3.6 Flash Google | 82.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.2 | 100% confidence 100 percent, Full |
| 101 | Gemma 4 26B A4B IT Google | 82.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 82.2 | 64% confidence 64 percent, Medium |
| 102 | Claude Opus 4.5 Anthropic | 81.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Nov 24, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.7 | 100% confidence 100 percent, Full |
| 103 | o4-mini OpenAI | 81.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Apr 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.7 | 87% confidence 87 percent, High |
| 104 | GPT-5 Nano OpenAI | 81.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Oct 31, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 81.1 | 85% confidence 85 percent, High |
| 105 | Qwen3.7 Plus Alibaba / Qwen | 80%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.0 | 64% confidence 64 percent, Medium |
| 106 | Claude Sonnet 5 Anthropic | 80%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.0 | 93% confidence 93 percent, High |
| 107 | Gemini 3.1 Flash Lite Google | 80%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.0 | 32% confidence 32 percent, Low |
| 108 | Gemini 3.5 Flash Google | 80%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.0 | 100% confidence 100 percent, Full |
| 109 | Gemini 3.6 Flash Google | 80%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.0 | 100% confidence 100 percent, Full |
| 110 | GLM-5 Z.ai | 80%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 12, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 80.0 | 90% confidence 90 percent, High |
| 111 | GPT-5.2 OpenAI | 78.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Dec 11, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.9 | 100% confidence 100 percent, Full |
| 112 | GPT-5 Mini OpenAI | 78.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 78.3 | 99% confidence 99 percent, High |
| 113 | Qwen3.7 Flash Alibaba / Qwen | 77.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 77.8 | 32% confidence 32 percent, Low |
| 114 | Claude Sonnet 4.5 Anthropic | 77.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 21, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 77.8 | 80% confidence 80 percent, High |
| 115 | Claude Sonnet 4.5 Anthropic | 77.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 59KPublished Oct 28, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 77.8 | 80% confidence 80 percent, High |
| 116 | Claude Sonnet 4.6 Anthropic | 75.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 75.6 | 100% confidence 100 percent, Full |
| 117 | GLM-5.2 Z.ai | 75.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 75.6 | 100% confidence 100 percent, Full |
| 118 | GPT-5 Nano OpenAI | 74.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 74.2 | 85% confidence 85 percent, High |
| 119 | Qwen3 Max Alibaba / Qwen | 73.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 6, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.3 | 69% confidence 69 percent, Medium |
| 120 | Gemma 4 31B IT Google | 73.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.3 | 64% confidence 64 percent, Medium |
| 121 | o4-mini OpenAI | 73.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 73.3 | 87% confidence 87 percent, High |
| 122 | Claude Sonnet 4 Anthropic | 71.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.1 | 100% confidence 100 percent, Full |
| 123 | Claude Sonnet 4.5 Anthropic | 71.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Oct 28, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.1 | 80% confidence 80 percent, High |
| 124 | Claude Sonnet 4.6 Anthropic | 71.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.1 | 100% confidence 100 percent, Full |
| 125 | Gemini 3.5 Flash Lite Google | 71.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.1 | 100% confidence 100 percent, Full |
| 126 | MiniMax-M3 MiniMax | 71.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 71.1 | 100% confidence 100 percent, Full |
| 127 | Qwen3.5 35B-A3B Alibaba / Qwen | 70%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 70.0 | 64% confidence 64 percent, Medium |
| 128 | Qwen3.6 35B-A3B Alibaba / Qwen | 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.9 | 37% confidence 37 percent, Low |
| 129 | Claude Opus 4.1 Anthropic | 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 27KPublished Aug 5, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.9 | 80% confidence 80 percent, High |
| 130 | Claude Sonnet 4 Anthropic | 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 59KPublished May 23, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.9 | 100% confidence 100 percent, Full |
| 131 | Kimi K3 Moonshot AI | 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.9 | 100% confidence 100 percent, Full |
| 132 | GPT-5.4 nano OpenAI | 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.9 | 100% confidence 100 percent, Full |
| 133 | GPT-5.6 Sol OpenAI | 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.9 | 100% confidence 100 percent, Full |
| 134 | GPT-5.5 Instant OpenAI | 68.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 2, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 68.1 | 64% confidence 64 percent, Medium |
| 135 | Qwen3 32B Alibaba / Qwen | 66.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.9 | 67% confidence 67 percent, Medium |
| 136 | Qwen3.6 27B Alibaba / Qwen | 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.7 | 53% confidence 53 percent, Medium |
| 137 | Claude Haiku 4.5 Anthropic | 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.7 | 64% confidence 64 percent, Medium |
| 138 | GPT-5.6 Luna OpenAI | 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.7 | 100% confidence 100 percent, Full |
| 139 | DeepSeek-R1 DeepSeek | 66.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 29, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 66.4 | 88% confidence 88 percent, High |
| 140 | GPT OSS 20B OpenAI | 65.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] mediumPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 65.3 | 64% confidence 64 percent, Medium |
| 141 | Claude Opus 4.1 Anthropic | 64.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Aug 5, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 64.4 | 80% confidence 80 percent, High |
| 142 | Claude Opus 4 Anthropic | 64.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 27KPublished May 28, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 64.4 | 100% confidence 100 percent, Full |
| 143 | GPT-5.1 OpenAI | 63.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Nov 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 63.9 | 85% confidence 85 percent, High |
| 144 | Qwen3 30B A3B Alibaba / Qwen | 62.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 62.8 | 64% confidence 64 percent, Medium |
| 145 | GPT-5.2 OpenAI | 62.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 62.2 | 100% confidence 100 percent, Full |
| 146 | Qwen3.5 9B Alibaba / Qwen | 61.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 61.7 | 32% confidence 32 percent, Low |
| 147 | Claude Opus 4 Anthropic | 60%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 60.0 | 100% confidence 100 percent, Full |
| 148 | Gemini 3.5 Flash Lite Google | 60%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 60.0 | 100% confidence 100 percent, Full |
| 149 | o3 OpenAI | 60%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 60.0 | 87% confidence 87 percent, High |
| 150 | QwQ 32B Alibaba / Qwen | 59.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 59.2 | 67% confidence 67 percent, Medium |
| 151 | GLM-4.7-Flash Z.ai | 58.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 58.3 | 69% confidence 69 percent, Medium |
| 152 | Claude Sonnet 3.7 Anthropic | 57.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Mar 13, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 57.8 | 100% confidence 100 percent, Full |
| 153 | GPT-5.4 OpenAI | 57.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Jul 15, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 57.8 | 100% confidence 100 percent, Full |
| 154 | GPT-5.5 OpenAI | 57.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 57.8 | 100% confidence 100 percent, Full |
| 155 | o4-mini OpenAI | 57.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 57.8 | 87% confidence 87 percent, High |
| 156 | DeepSeek-R1-Distill-Qwen-32B DeepSeek | 55.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 55.6 | 32% confidence 32 percent, Low |
| 157 | GPT-5 Mini OpenAI | 55.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 55.6 | 99% confidence 99 percent, High |
| 158 | Qwen3.5 35B-A3B Alibaba / Qwen | 54.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 54.4 | 64% confidence 64 percent, Medium |
| 159 | Claude Sonnet 3.7 Anthropic | 53.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Mar 12, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 53.3 | 100% confidence 100 percent, Full |
| 160 | Claude Sonnet 4 Anthropic | 53.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 53.3 | 100% confidence 100 percent, Full |
| 161 | GPT-5.6 Terra OpenAI | 53.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 53.3 | 100% confidence 100 percent, Full |
| 162 | Gemini 3.5 Flash Lite Google | 51.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 51.1 | 100% confidence 100 percent, Full |
| 163 | GPT OSS 20B OpenAI | 50.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 50.8 | 64% confidence 64 percent, Medium |
| 164 | DeepSeek Chat DeepSeek | 48.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 16, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 48.9 | 32% confidence 32 percent, Low |
| 165 | Claude Opus 4.5 Anthropic | 48.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Nov 24, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 48.1 | 100% confidence 100 percent, Full |
| 166 | Claude Sonnet 3.7 Anthropic | 46.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Feb 26, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.7 | 100% confidence 100 percent, Full |
| 167 | DeepSeek V4 Pro DeepSeek | 46.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.7 | 85% confidence 85 percent, High |
| 168 | GPT-5 OpenAI | 46.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 20, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.7 | 100% confidence 100 percent, Full |
| 169 | GPT-5 Nano OpenAI | 46.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.7 | 85% confidence 85 percent, High |
| 170 | GPT-5.4 nano OpenAI | 46.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 46.7 | 100% confidence 100 percent, Full |
| 171 | GPT-4.1 mini OpenAI | 44.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 44.7 | 87% confidence 87 percent, High |
| 172 | Gemini 3.1 Flash Lite Google | 44.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 44.4 | 32% confidence 32 percent, Low |
| 173 | Qwen3.5 9B Alibaba / Qwen | 42.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 42.8 | 32% confidence 32 percent, Low |
| 174 | Claude Opus 4 Anthropic | 42.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 42.2 | 100% confidence 100 percent, Full |
| 175 | GPT OSS 20B OpenAI | 40.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 40.3 | 64% confidence 64 percent, Medium |
| 176 | Claude Opus 4.1 Anthropic | 40%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 5, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 40.0 | 80% confidence 80 percent, High |
| 177 | GPT-5.6 Luna OpenAI | 40%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 40.0 | 100% confidence 100 percent, Full |
| 178 | GPT-4.1 OpenAI | 38.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 38.3 | 100% confidence 100 percent, Full |
| 179 | DeepSeek V3 0324 DeepSeek | 37.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 1, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 37.8 | 74% confidence 74 percent, Medium |
| 180 | Gemini 3.1 Flash Lite Google | 37.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 37.8 | 32% confidence 32 percent, Low |
| 181 | GPT-5.1 OpenAI | 37.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 37.8 | 85% confidence 85 percent, High |
| 182 | Claude Haiku 4.5 Anthropic | 35.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 16, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 35.8 | 64% confidence 64 percent, Medium |
| 183 | Claude Sonnet 4.5 Anthropic | 35.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Sep 29, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 35.6 | 80% confidence 80 percent, High |
| 184 | GPT-5 Nano OpenAI | 35.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] minimalPublished Jul 13, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 35.6 | 85% confidence 85 percent, High |
| 185 | Mistral Medium 3 Mistral AI | 32.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 32.2 | 85% confidence 85 percent, High |
| 186 | Magistral Small Mistral AI | 30%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 30.0 | 48% confidence 48 percent, Low |
| 187 | Claude Sonnet 4 Anthropic | 28.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 28.9 | 100% confidence 100 percent, Full |
| 188 | GPT-4.1 nano OpenAI | 28.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 14, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 28.9 | 82% confidence 82 percent, High |
| 189 | GLM-5.2 Z.ai | 28.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 28.9 | 100% confidence 100 percent, Full |
| 190 | Magistral Small 1.2 Mistral AI | 28.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 28.1 | 32% confidence 32 percent, Low |
| 191 | MiniMax-M3 MiniMax | 26.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 10, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 26.7 | 100% confidence 100 percent, Full |
| 192 | GPT-5.4 mini OpenAI | 26.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 7, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 26.7 | 100% confidence 100 percent, Full |
| 193 | Qwen3 30B A3B Alibaba / Qwen | 25.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 25.6 | 64% confidence 64 percent, Medium |
| 194 | GLM-4.7-Flash Z.ai | 25%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 25.0 | 69% confidence 69 percent, Medium |
| 195 | Qwen3 32B Alibaba / Qwen | 23.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] nonePublished Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 23.1 | 67% confidence 67 percent, Medium |
| 196 | Gemma 3 27B IT Google | 22.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 22.5 | 74% confidence 74 percent, Medium |
| 197 | Claude Sonnet 3.7 Anthropic | 21.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 21.9 | 100% confidence 100 percent, Full |
| 198 | Llama 4 Maverick 17B Instruct Meta | 20.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 20.6 | 78% confidence 78 percent, Medium |
| 199 | Gemma 3 12B IT Google | 16.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 16.7 | 64% confidence 64 percent, Medium |
| 200 | DeepSeek-V3 DeepSeek | 15.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 15.8 | 93% confidence 93 percent, High |
| 201 | Claude Sonnet 3.5 v2 Anthropic | 8.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 8.5 | 100% confidence 100 percent, Full |
| 202 | Llama 4 Scout 17B Instruct Meta | 7.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 8, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 7.8 | 64% confidence 64 percent, Medium |
| 203 | Mistral Large 2.1 Mistral AI | 7.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 7.8 | 100% confidence 100 percent, Full |
| 204 | Gemma 3 4B IT Google | 7.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 7.5 | 64% confidence 64 percent, Medium |
| 205 | GPT-4o mini OpenAI | 6.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jul 30, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 6.9 | 100% confidence 100 percent, Full |
| 206 | GPT-4o (2024-08-06) OpenAI | 6.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 6.4 | 100% confidence 100 percent, Full |
| 207 | GPT-4o (2024-05-13) OpenAI | 6.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 6.3 | 93% confidence 93 percent, High |
| 208 | GPT-4o (2024-11-20) OpenAI | 6.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 6.3 | 61% confidence 61 percent, Medium |
| 209 | Qwen Turbo Alibaba / Qwen | 6.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Apr 7, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 6.1 | 47% confidence 47 percent, Low |
| 210 | Mistral Small 3.1 24B Mistral AI | 5.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Mar 18, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 5.8 | 64% confidence 64 percent, Medium |
| 211 | Llama-3.3-70B-Instruct Meta | 5.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 5.1 | 100% confidence 100 percent, Full |
| 212 | Claude Haiku 3.5 Anthropic | 4.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 4.3 | 100% confidence 100 percent, Full |
| 213 | Llama-3.1-70B-Instruct Meta | 3.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 3.6 | 93% confidence 93 percent, High |
| 214 | Claude Haiku 3 Anthropic | 1.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 1.8 | 93% confidence 93 percent, High |
| 215 | Llama-3.1-8B-Instruct Meta | 1.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 27, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 1.7 | 93% confidence 93 percent, High |
| 216 | Llama-3.2-1B Meta | 0.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026
Retrieved Oct 9, 2026 · CC-BY Open source ↗ | 0.6 | 93% confidence 93 percent, High |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the math pillar of the SI Score.