ARC-AGI-1 (public eval)Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 Anthropic | 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-1-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 99.0 | 100% confidence 100 percent, Full |
| 2 | Claude Opus 5 Anthropic | 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 99.0 | 100% confidence 100 percent, Full |
| 3 | Claude Opus 5 Anthropic | 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 99.0 | 100% confidence 100 percent, Full |
| 4 | GPT-5.6 Sol OpenAI | 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 99.0 | 100% confidence 100 percent, Full |
| 5 | GPT-6 Astra OpenAI | 99%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 99.0 | 100% confidence 100 percent, Full |
| 6 | Claude Fable 5.1 Anthropic | 98.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-1-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.8 | 100% confidence 100 percent, Full |
| 7 | GPT-5.6 Sol OpenAI | 98.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.8 | 100% confidence 100 percent, Full |
| 8 | GPT-6 Astra OpenAI | 98.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.8 | 100% confidence 100 percent, Full |
| 9 | Claude Opus 5.5 Anthropic | 98.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.6 | 100% confidence 100 percent, Full |
| 10 | GPT-6 Astra OpenAI | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 11 | GPT-6.1 Sol OpenAI | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 12 | GPT-6.1 Sol OpenAI | 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.5 | 100% confidence 100 percent, Full |
| 13 | Claude Opus 5.5 Anthropic | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.3 | 100% confidence 100 percent, Full |
| 14 | Claude Opus 5.5 Anthropic | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.3 | 100% confidence 100 percent, Full |
| 15 | Claude Opus 5.5 Anthropic | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.3 | 100% confidence 100 percent, Full |
| 16 | GPT-5.4 Pro OpenAI | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-pro-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.3 | 69% confidence 69 percent, Medium |
| 17 | GPT-5.6 Terra OpenAI | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.3 | 100% confidence 100 percent, Full |
| 18 | GPT-6 Astra OpenAI | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.3 | 100% confidence 100 percent, Full |
| 19 | GPT-6.1 Sol OpenAI | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.3 | 100% confidence 100 percent, Full |
| 20 | Claude Fable 5 Anthropic | 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.0 | 100% confidence 100 percent, Full |
| 21 | Claude Fable 5 Anthropic | 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.0 | 100% confidence 100 percent, Full |
| 22 | DeepSeek V4.1 Flash DeepSeek | 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.0 | 69% confidence 69 percent, Medium |
| 23 | Gemini 3.8 Flash Google | 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.0 | 100% confidence 100 percent, Full |
| 24 | GPT-5.5 Pro OpenAI | 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.0 | 53% confidence 53 percent, Medium |
| 25 | GPT-5.5 Pro OpenAI | 98.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.0 | 53% confidence 53 percent, Medium |
| 26 | Claude Fable 5 Anthropic | 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.9 | 100% confidence 100 percent, Full |
| 27 | GPT-6.1 Sol OpenAI | 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.9 | 100% confidence 100 percent, Full |
| 28 | Gemini 3.8 Flash Google | 97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 29 | GPT-6 Astra OpenAI | 97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 30 | GPT-6 Sol OpenAI | 97.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.8 | 100% confidence 100 percent, Full |
| 31 | Claude Fable 5 Anthropic | 97.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.6 | 100% confidence 100 percent, Full |
| 32 | Claude Fable 5.1 Anthropic | 97.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-1-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.6 | 100% confidence 100 percent, Full |
| 33 | GPT-5.2 Pro OpenAI | 97.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.6 | 48% confidence 48 percent, Low |
| 34 | GPT-5.5 OpenAI | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 35 | GPT-5.5 OpenAI | 97.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.4 | 100% confidence 100 percent, Full |
| 36 | GPT-5.6 Sol OpenAI | 97.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.3 | 100% confidence 100 percent, Full |
| 37 | Gemini 3.1 Pro Preview Google | 97.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-1-pro-previewPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.2 | 100% confidence 100 percent, Full |
| 38 | Claude Opus 4.7 Anthropic | 97%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.0 | 100% confidence 100 percent, Full |
| 39 | Claude Opus 4.6 Anthropic | 96.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude-opus-4-6-thinking-120K-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.8 | 100% confidence 100 percent, Full |
| 40 | Claude Opus 4.7 Anthropic | 96.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.6 | 100% confidence 100 percent, Full |
| 41 | Claude Fable 5.1 Anthropic | 96.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-1-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.6 | 100% confidence 100 percent, Full |
| 42 | Gemini 3.7 Flash Google | 96.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] google-gemini-3-7-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.6 | 100% confidence 100 percent, Full |
| 43 | Claude Fable 5 Anthropic | 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.5 | 100% confidence 100 percent, Full |
| 44 | GPT-5.5 OpenAI | 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.5 | 100% confidence 100 percent, Full |
| 45 | GPT-6 Sol OpenAI | 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.5 | 100% confidence 100 percent, Full |
| 46 | GPT-5.4 OpenAI | 96.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.4 | 100% confidence 100 percent, Full |
| 47 | GPT-6.1 Sol OpenAI | 96.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.4 | 100% confidence 100 percent, Full |
| 48 | Claude Opus 4.6 Anthropic | 96.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude-opus-4-6-thinking-120K-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.3 | 100% confidence 100 percent, Full |
| 49 | DeepSeek V4.1 Flash DeepSeek | 96.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] deepseek-v4-1-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.3 | 69% confidence 69 percent, Medium |
| 50 | Grok 4.6 xAI | 96.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] xai-grok-4-6-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.3 | 100% confidence 100 percent, Full |
| 51 | Grok 4.6 xAI | 96.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] xai-grok-4-6-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.3 | 100% confidence 100 percent, Full |
| 52 | Claude Opus 4.7 Anthropic | 96.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.1 | 100% confidence 100 percent, Full |
| 53 | Gemini 3.5 Flash Google | 96%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.0 | 100% confidence 100 percent, Full |
| 54 | Gemini 3.6 Flash Google | 96%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-6-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.0 | 100% confidence 100 percent, Full |
| 55 | DeepSeek V4.1 Flash DeepSeek | 95.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.9 | 69% confidence 69 percent, Medium |
| 56 | Claude Sonnet 4.6 Anthropic | 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude_sonnet_4_6_maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.8 | 100% confidence 100 percent, Full |
| 57 | GPT-5.6 Terra OpenAI | 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.8 | 100% confidence 100 percent, Full |
| 58 | Grok 4.7 xAI | 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.8 | 100% confidence 100 percent, Full |
| 59 | GPT-5.6 Terra OpenAI | 95.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.6 | 100% confidence 100 percent, Full |
| 60 | GPT-5.4 OpenAI | 95.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.6 | 100% confidence 100 percent, Full |
| 61 | Claude Fable 5.1 Anthropic | 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-1-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.5 | 100% confidence 100 percent, Full |
| 62 | DeepSeek V4 Pro 0813 DeepSeek | 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-pro-0813-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.5 | 69% confidence 69 percent, Medium |
| 63 | Grok 4.20 (Reasoning) xAI | 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] grok-4.20-beta-0309b-reasoningPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.5 | 80% confidence 80 percent, High |
| 64 | Claude Sonnet 4.6 Anthropic | 95.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude_sonnet_4_6_highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.3 | 100% confidence 100 percent, Full |
| 65 | Gemini 3.7 Flash Google | 95.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] google-gemini-3-7-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.3 | 100% confidence 100 percent, Full |
| 66 | Grok 4.7 xAI | 95.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.1 | 100% confidence 100 percent, Full |
| 67 | GPT-5.2 OpenAI | 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.0 | 100% confidence 100 percent, Full |
| 68 | Grok 4.7 xAI | 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.0 | 100% confidence 100 percent, Full |
| 69 | Kimi K3 Moonshot AI | 94.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.9 | 100% confidence 100 percent, Full |
| 70 | Claude Opus 4.6 Anthropic | 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'medium' output effort. [variant] claude-opus-4-6-thinking-120K-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.8 | 100% confidence 100 percent, Full |
| 71 | DeepSeek V4 Flash 0731 DeepSeek | 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.8 | 69% confidence 69 percent, Medium |
| 72 | Grok 4.5 xAI | 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.8 | 100% confidence 100 percent, Full |
| 73 | Grok 4.6 xAI | 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] xai-grok-4-6-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.8 | 100% confidence 100 percent, Full |
| 74 | GPT-5.2 Pro OpenAI | 94.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.6 | 48% confidence 48 percent, Low |
| 75 | Claude Opus 4.7 Anthropic | 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.5 | 100% confidence 100 percent, Full |
| 76 | Gemini 3.8 Flash Google | 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.5 | 100% confidence 100 percent, Full |
| 77 | GPT-5.6 Sol OpenAI | 93.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.9 | 100% confidence 100 percent, Full |
| 78 | DeepSeek V4 Pro 0813 DeepSeek | 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-pro-0813-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.3 | 69% confidence 69 percent, Medium |
| 79 | DeepSeek V4 Flash 0731 DeepSeek | 93%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.0 | 69% confidence 69 percent, Medium |
| 80 | Kimi K3 Moonshot AI | 92.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.9 | 100% confidence 100 percent, Full |
| 81 | DeepSeek V4 Pro 0813 DeepSeek | 92.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-pro-0813-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.8 | 69% confidence 69 percent, Medium |
| 82 | Gemini 3.6 Flash Google | 92.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-6-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.8 | 100% confidence 100 percent, Full |
| 83 | DeepSeek V4 Flash 0731 DeepSeek | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 69% confidence 69 percent, Medium |
| 84 | GPT-6 Luna OpenAI | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 85 | Grok 4.5 xAI | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 86 | GLM-5.3-Flash Z.ai | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] zai-glm-5-3-flash-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 87 | Claude Opus 5.5 Anthropic | 92.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.3 | 100% confidence 100 percent, Full |
| 88 | GPT-5.4 OpenAI | 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.0 | 100% confidence 100 percent, Full |
| 89 | GPT-6 Sol OpenAI | 91.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.4 | 100% confidence 100 percent, Full |
| 90 | Gemini 3.7 Flash Google | 91.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] google-gemini-3-7-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.3 | 100% confidence 100 percent, Full |
| 91 | GPT-5.6 Luna OpenAI | 90.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.8 | 100% confidence 100 percent, Full |
| 92 | Inkling Thinking Machines Lab | 90.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] thinky-inklingPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.6 | 100% confidence 100 percent, Full |
| 93 | GPT-5.2 OpenAI | 90.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.3 | 100% confidence 100 percent, Full |
| 94 | GPT-5.2 Pro OpenAI | 90.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.3 | 48% confidence 48 percent, Low |
| 95 | GPT-5.6 Luna OpenAI | 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 96 | Claude Opus 4.6 Anthropic | 89.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'low' output effort. [variant] claude-opus-4-6-thinking-120K-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 89.6 | 100% confidence 100 percent, Full |
| 97 | Inkling Small Thinking Machines Lab | 89.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] thinky-inkling-small-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 89.4 | 100% confidence 100 percent, Full |
| 98 | GPT-6 Sol OpenAI | 88.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.9 | 100% confidence 100 percent, Full |
| 99 | Gemini 3 Flash Preview Google | 88.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.3 | 93% confidence 93 percent, High |
| 100 | Qwen3.8 27B Alibaba / Qwen | 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] alibaba-qwen3-8-27b-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.5 | 69% confidence 69 percent, Medium |
| 101 | Inkling Small Thinking Machines Lab | 86.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] thinky-inkling-small-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.6 | 100% confidence 100 percent, Full |
| 102 | Claude Opus 4.5 Anthropic | 86.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.6 | 100% confidence 100 percent, Full |
| 103 | Grok 4.5 xAI | 86%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.0 | 100% confidence 100 percent, Full |
| 104 | GPT-6 Luna OpenAI | 85.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 85.8 | 100% confidence 100 percent, Full |
| 105 | GPT-5.6 Terra OpenAI | 84.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.8 | 100% confidence 100 percent, Full |
| 106 | GPT-5.6 Sol OpenAI | 84.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.5 | 100% confidence 100 percent, Full |
| 107 | Grok 4.6 xAI | 84%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] xai-grok-4-6-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.0 | 100% confidence 100 percent, Full |
| 108 | GPT-5.5 OpenAI | 82.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 82.3 | 100% confidence 100 percent, Full |
| 109 | Grok 4.7 xAI | 82.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] xai-grok-4-7-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 82.3 | 100% confidence 100 percent, Full |
| 110 | Claude Opus 4.5 Anthropic | 81.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 81.6 | 100% confidence 100 percent, Full |
| 111 | Gemini 3.6 Flash Google | 81.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-6-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 81.4 | 100% confidence 100 percent, Full |
| 112 | GPT-5.2 OpenAI | 80.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 80.6 | 100% confidence 100 percent, Full |
| 113 | GLM-5.2 Z.ai | 80.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5.2Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 114 | GPT-5.4 OpenAI | 80%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 80.0 | 100% confidence 100 percent, Full |
| 115 | GPT-6 Luna OpenAI | 79.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 79.8 | 100% confidence 100 percent, Full |
| 116 | GPT-5.6 Luna OpenAI | 79.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 79.3 | 100% confidence 100 percent, Full |
| 117 | GPT-6 Sol OpenAI | 79.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 79.3 | 100% confidence 100 percent, Full |
| 118 | GLM-5.3-Flash Z.ai | 78.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] zai-glm-5-3-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 78.6 | 100% confidence 100 percent, Full |
| 119 | Inkling Small Thinking Machines Lab | 77.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] thinky-inkling-small-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 77.8 | 100% confidence 100 percent, Full |
| 120 | GPT-5.1 OpenAI | 77.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 77.1 | 85% confidence 85 percent, High |
| 121 | GPT-5 Pro OpenAI | 77%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-pro-2025-10-06Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 77.0 | 64% confidence 64 percent, Medium |
| 122 | Qwen3.8 27B Alibaba / Qwen | 76.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] alibaba-qwen3-8-27b-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 76.3 | 69% confidence 69 percent, Medium |
| 123 | Qwen3.8 27B Alibaba / Qwen | 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] alibaba-qwen3-8-27b-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 75.8 | 69% confidence 69 percent, Medium |
| 124 | Kimi K3 Moonshot AI | 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 75.8 | 100% confidence 100 percent, Full |
| 125 | GPT-5.4 mini OpenAI | 75.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 75.1 | 100% confidence 100 percent, Full |
| 126 | Claude Sonnet 4.5 (latest) Anthropic | 73.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 73.8 | 43% confidence 43 percent, Low |
| 127 | Kimi K2.5 Moonshot AI | 73.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] kimi-k2.5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 73.1 | 100% confidence 100 percent, Full |
| 128 | GPT-5.6 Terra OpenAI | 70.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 70.8 | 100% confidence 100 percent, Full |
| 129 | GPT-6 Luna OpenAI | 70.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 70.1 | 100% confidence 100 percent, Full |
| 130 | Claude Opus 4.5 Anthropic | 70.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 70.1 | 100% confidence 100 percent, Full |
| 131 | GPT-5.1 OpenAI | 68.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.9 | 85% confidence 85 percent, High |
| 132 | o4-mini OpenAI | 68.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.0 | 87% confidence 87 percent, High |
| 133 | Gemini 3 Flash Preview Google | 67.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.9 | 93% confidence 93 percent, High |
| 134 | Gemini 3.5 Flash Lite Google | 66.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-5-flash-lite-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 66.4 | 100% confidence 100 percent, Full |
| 135 | GPT-5.4 mini OpenAI | 66.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 66.3 | 100% confidence 100 percent, Full |
| 136 | Inkling Small Thinking Machines Lab | 66.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] thinky-inkling-small-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 66.1 | 100% confidence 100 percent, Full |
| 137 | GPT-5.2 OpenAI | 65.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.9 | 100% confidence 100 percent, Full |
| 138 | GPT-5 OpenAI | 65.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.9 | 100% confidence 100 percent, Full |
| 139 | GPT-5.6 Luna OpenAI | 64.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 64.4 | 100% confidence 100 percent, Full |
| 140 | o3 OpenAI | 64.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 64.3 | 87% confidence 87 percent, High |
| 141 | Claude Sonnet 4.5 (latest) Anthropic | 63.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.6 | 43% confidence 43 percent, Low |
| 142 | GPT-5 OpenAI | 63.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.4 | 100% confidence 100 percent, Full |
| 143 | o3-pro OpenAI | 63.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.3 | 22% confidence 22 percent, Low |
| 144 | Claude Haiku 4.5 (latest) Anthropic | 62.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 62.9 | 33% confidence 33 percent, Low |
| 145 | DeepSeek V3.2 DeepSeek | 61.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek-v3.2Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 61.6 | 56% confidence 56 percent, Medium |
| 146 | GPT-5 Mini OpenAI | 61.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 61.5 | 99% confidence 99 percent, High |
| 147 | MiniMax-M2.5 MiniMax | 59.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] minimax-m2.5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.1 | 100% confidence 100 percent, Full |
| 148 | GLM-5 Z.ai | 58.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.6 | 90% confidence 90 percent, High |
| 149 | o3-pro OpenAI | 58.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.1 | 22% confidence 22 percent, Low |
| 150 | GLM-5.3-Flash Z.ai | 57.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] zai-glm-5-3-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 57.8 | 100% confidence 100 percent, Full |
| 151 | Claude Sonnet 4 Anthropic | 56.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 56.8 | 100% confidence 100 percent, Full |
| 152 | o3 OpenAI | 56.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 56.7 | 87% confidence 87 percent, High |
| 153 | Gemini 2.5 Pro Google | 56.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 56.4 | 90% confidence 90 percent, High |
| 154 | Gemini 2.5 Pro Google | 55.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 55.9 | 90% confidence 90 percent, High |
| 155 | GPT-5.4 mini OpenAI | 55.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 55.4 | 100% confidence 100 percent, Full |
| 156 | Claude Opus 4 Anthropic | 54.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 54.3 | 100% confidence 100 percent, Full |
| 157 | Claude Sonnet 4.5 (latest) Anthropic | 53.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 53.5 | 43% confidence 43 percent, Low |
| 158 | Claude Opus 4.5 Anthropic | 52.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 52.6 | 100% confidence 100 percent, Full |
| 159 | GPT-5.4 nano OpenAI | 51.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 51.6 | 100% confidence 100 percent, Full |
| 160 | Claude Haiku 4.5 (latest) Anthropic | 51.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 51.4 | 33% confidence 33 percent, Low |
| 161 | o3-pro OpenAI | 50.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 50.9 | 22% confidence 22 percent, Low |
| 162 | o4-mini OpenAI | 50.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 50.2 | 87% confidence 87 percent, High |
| 163 | Claude Sonnet 4 Anthropic | 48.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 48.6 | 100% confidence 100 percent, Full |
| 164 | GPT-5 OpenAI | 48.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 48.4 | 100% confidence 100 percent, Full |
| 165 | GPT-5.4 nano OpenAI | 47.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 47.9 | 100% confidence 100 percent, Full |
| 166 | o3 OpenAI | 47.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 47.6 | 87% confidence 87 percent, High |
| 167 | GPT-5.6 Luna OpenAI | 47.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 47.4 | 100% confidence 100 percent, Full |
| 168 | Gemini 3.5 Flash Lite Google | 47%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-5-flash-lite-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 47.0 | 100% confidence 100 percent, Full |
| 169 | o3-mini OpenAI | 46.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 46.6 | 96% confidence 96 percent, High |
| 170 | GPT-6 Luna OpenAI | 46.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 46.5 | 100% confidence 100 percent, Full |
| 171 | GPT-5 Mini OpenAI | 46.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 46.3 | 99% confidence 99 percent, High |
| 172 | Claude Opus 4 Anthropic | 45.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 45.6 | 100% confidence 100 percent, Full |
| 173 | Claude Haiku 4.5 (latest) Anthropic | 45%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 45.0 | 33% confidence 33 percent, Low |
| 174 | Gemini 2.5 Pro Google | 44.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 44.2 | 90% confidence 90 percent, High |
| 175 | GPT-5.1 OpenAI | 44%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 44.0 | 85% confidence 85 percent, High |
| 176 | GPT-5.4 nano OpenAI | 43.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 43.4 | 100% confidence 100 percent, Full |
| 177 | Claude Opus 4 Anthropic | 43.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 43.3 | 100% confidence 100 percent, Full |
| 178 | Gemini 3 Flash Preview Google | 38.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 38.2 | 93% confidence 93 percent, High |
| 179 | Claude Sonnet 4.5 (latest) Anthropic | 36.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 36.6 | 43% confidence 43 percent, Low |
| 180 | Claude Opus 4 Anthropic | 35.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 35.5 | 100% confidence 100 percent, Full |
| 181 | Claude Sonnet 4.5 (latest) Anthropic | 35.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 35.4 | 43% confidence 43 percent, Low |
| 182 | Claude Sonnet 4 Anthropic | 33%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 33.0 | 100% confidence 100 percent, Full |
| 183 | GPT-5.4 mini OpenAI | 31.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 31.8 | 100% confidence 100 percent, Full |
| 184 | Claude Sonnet 4 Anthropic | 31.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 31.3 | 100% confidence 100 percent, Full |
| 185 | o3-mini OpenAI | 30.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 30.6 | 96% confidence 96 percent, High |
| 186 | GPT-5 Nano OpenAI | 29.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 29.7 | 85% confidence 85 percent, High |
| 187 | o4-mini OpenAI | 27.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 27.6 | 87% confidence 87 percent, High |
| 188 | Claude Haiku 4.5 (latest) Anthropic | 27.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 27.1 | 33% confidence 33 percent, Low |
| 189 | DeepSeek-R1 DeepSeek | 27.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 27.0 | 88% confidence 88 percent, High |
| 190 | Claude Haiku 4.5 (latest) Anthropic | 26.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 26.6 | 33% confidence 33 percent, Low |
| 191 | Gemini 3.5 Flash Lite Google | 24.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-5-flash-lite-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 24.6 | 100% confidence 100 percent, Full |
| 192 | GPT-5.4 nano OpenAI | 24.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 24.6 | 100% confidence 100 percent, Full |
| 193 | GPT-5 Mini OpenAI | 24.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 24.4 | 99% confidence 99 percent, High |
| 194 | GPT-5 Nano OpenAI | 20.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 20.8 | 85% confidence 85 percent, High |
| 195 | Gemini 2.5 Pro Google | 17.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 17.5 | 90% confidence 90 percent, High |
| 196 | o3-mini OpenAI | 17.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 17.4 | 96% confidence 96 percent, High |
| 197 | Qwen3 235B-A22B Instruct 2507 Alibaba / Qwen | 17%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] qwen3-235b-a22b-instruct-2507Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 17.0 | 48% confidence 48 percent, Low |
| 198 | GPT-5.2 OpenAI | 16.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 16.5 | 100% confidence 100 percent, Full |
| 199 | GPT-5.1 OpenAI | 12.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 12.4 | 85% confidence 85 percent, High |
| 200 | GPT-5 Nano OpenAI | 11.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 11.8 | 85% confidence 85 percent, High |
| 201 | GPT-4.1 OpenAI | 11.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 11.8 | 100% confidence 100 percent, Full |
| 202 | Magistral Medium (latest) Mistral AI | 8.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 8.9 | 16% confidence 16 percent, Low |
| 203 | Magistral Small Mistral AI | 8.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 8.6 | 48% confidence 48 percent, Low |
| 204 | Magistral Medium (latest) Mistral AI | 8.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506-thinkingPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 8.0 | 16% confidence 16 percent, Low |
| 205 | GPT-4.1 mini OpenAI | 7.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-mini-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.2 | 87% confidence 87 percent, High |
| 206 | Llama 4 Maverick 17B Instruct Meta | 7.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Llama-4-Maverick-17B-128E-Instruct-FP8-togetherPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.1 | 78% confidence 78 percent, Medium |
| 207 | GPT-4.1 nano OpenAI | 1.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-nano-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.8 | 82% confidence 82 percent, High |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.