ARC-AGI-2 (semi-private)Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | GPT-6 Astra OpenAI | 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.0 | 100% confidence 100 percent, Full |
| 2 | GPT-6.1 Sol OpenAI | 94.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.2 | 100% confidence 100 percent, Full |
| 3 | Claude Opus 5.5 Anthropic | 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.3 | 100% confidence 100 percent, Full |
| 4 | GPT-6 Astra OpenAI | 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.3 | 100% confidence 100 percent, Full |
| 5 | Claude Opus 5.5 Anthropic | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 6 | GPT-5.6 Sol OpenAI | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 7 | GPT-6 Astra OpenAI | 92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.1 | 100% confidence 100 percent, Full |
| 8 | GPT-6 Astra OpenAI | 92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.1 | 100% confidence 100 percent, Full |
| 9 | Claude Opus 5.5 Anthropic | 91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.7 | 100% confidence 100 percent, Full |
| 10 | GPT-6.1 Sol OpenAI | 91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.7 | 100% confidence 100 percent, Full |
| 11 | GPT-6.1 Sol OpenAI | 91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.7 | 100% confidence 100 percent, Full |
| 12 | Claude Opus 5 Anthropic | 90.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.4 | 100% confidence 100 percent, Full |
| 13 | Claude Fable 5.1 Anthropic | 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-1-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 14 | Claude Fable 5.1 Anthropic | 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-1-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 15 | GPT-5.6 Sol OpenAI | 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 16 | GPT-6 Sol OpenAI | 89.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 89.6 | 100% confidence 100 percent, Full |
| 17 | Claude Fable 5 Anthropic | 89.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 89.2 | 100% confidence 100 percent, Full |
| 18 | Gemini 3.8 Flash Google | 89.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 89.2 | 100% confidence 100 percent, Full |
| 19 | Claude Fable 5.1 Anthropic | 88.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-1-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.8 | 100% confidence 100 percent, Full |
| 20 | Claude Fable 5 Anthropic | 88.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.3 | 100% confidence 100 percent, Full |
| 21 | Claude Opus 5 Anthropic | 88.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.3 | 100% confidence 100 percent, Full |
| 22 | Claude Fable 5 Anthropic | 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.5 | 100% confidence 100 percent, Full |
| 23 | Claude Opus 5.5 Anthropic | 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.5 | 100% confidence 100 percent, Full |
| 24 | GPT-6.1 Sol OpenAI | 86.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.7 | 100% confidence 100 percent, Full |
| 25 | Claude Fable 5.1 Anthropic | 86.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-1-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.3 | 100% confidence 100 percent, Full |
| 26 | GPT-5.6 Sol OpenAI | 85.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 85.4 | 100% confidence 100 percent, Full |
| 27 | GPT-6 Astra OpenAI | 85.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 85.4 | 100% confidence 100 percent, Full |
| 28 | GPT-5.5 OpenAI | 85%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 85.0 | 100% confidence 100 percent, Full |
| 29 | Gemini 3.7 Flash Google | 84.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] google-gemini-3-7-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.6 | 100% confidence 100 percent, Full |
| 30 | GPT-5.5 Pro OpenAI | 84.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.6 | 53% confidence 53 percent, Medium |
| 31 | GPT-5.5 Pro OpenAI | 84.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.2 | 53% confidence 53 percent, Medium |
| 32 | GPT-5.6 Terra OpenAI | 83.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 83.9 | 100% confidence 100 percent, Full |
| 33 | GPT-5.4 Pro OpenAI | 83.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-pro-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 83.3 | 69% confidence 69 percent, Medium |
| 34 | GPT-5.5 OpenAI | 83.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 83.3 | 100% confidence 100 percent, Full |
| 35 | Gemini 3.8 Flash Google | 82.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 82.9 | 100% confidence 100 percent, Full |
| 36 | Claude Fable 5 Anthropic | 82.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 82.5 | 100% confidence 100 percent, Full |
| 37 | Claude Fable 5.1 Anthropic | 78.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-1-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 78.3 | 100% confidence 100 percent, Full |
| 38 | GPT-6 Sol OpenAI | 78.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 78.1 | 100% confidence 100 percent, Full |
| 39 | Gemini 3.8 Flash Google | 77.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 77.5 | 100% confidence 100 percent, Full |
| 40 | Gemini 3.1 Pro Preview Google | 77.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-1-pro-previewPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 77.1 | 100% confidence 100 percent, Full |
| 41 | Claude Fable 5 Anthropic | 76.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 76.8 | 100% confidence 100 percent, Full |
| 42 | GPT-6.1 Sol OpenAI | 76.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 76.7 | 100% confidence 100 percent, Full |
| 43 | Claude Opus 4.7 Anthropic | 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 75.8 | 100% confidence 100 percent, Full |
| 44 | GPT-5.6 Terra OpenAI | 74.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 74.2 | 100% confidence 100 percent, Full |
| 45 | GPT-5.4 OpenAI | 74.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 74.0 | 100% confidence 100 percent, Full |
| 46 | DeepSeek V4.1 Flash DeepSeek | 72.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 72.9 | 69% confidence 69 percent, Medium |
| 47 | Claude Opus 4.8 Anthropic | 72.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-opus-4-8-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 72.1 | 100% confidence 100 percent, Full |
| 48 | Gemini 3.5 Flash Google | 72.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 72.1 | 100% confidence 100 percent, Full |
| 49 | Claude Opus 4.8 Anthropic | 71.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' output effort. [variant] anthropic-opus-4-8-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 71.7 | 100% confidence 100 percent, Full |
| 50 | GPT-5.5 OpenAI | 70.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 70.4 | 100% confidence 100 percent, Full |
| 51 | Claude Opus 5.5 Anthropic | 70.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 70.1 | 100% confidence 100 percent, Full |
| 52 | Claude Opus 4.6 Anthropic | 69.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude-opus-4-6-thinking-120K-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 69.2 | 100% confidence 100 percent, Full |
| 53 | GPT-6 Sol OpenAI | 68.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.9 | 100% confidence 100 percent, Full |
| 54 | Claude Opus 4.6 Anthropic | 68.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude-opus-4-6-thinking-120K-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.8 | 100% confidence 100 percent, Full |
| 55 | Claude Opus 4.7 Anthropic | 68.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.3 | 100% confidence 100 percent, Full |
| 56 | Claude Opus 4.7 Anthropic | 67.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.5 | 100% confidence 100 percent, Full |
| 57 | DeepSeek V4.1 Flash DeepSeek | 67.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.5 | 69% confidence 69 percent, Medium |
| 58 | GPT-5.4 OpenAI | 67.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.5 | 100% confidence 100 percent, Full |
| 59 | Grok 4.6 xAI | 67.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] xai-grok-4-6-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.1 | 100% confidence 100 percent, Full |
| 60 | GPT-5.6 Sol OpenAI | 67.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.1 | 100% confidence 100 percent, Full |
| 61 | GPT-5.6 Terra OpenAI | 67.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.1 | 100% confidence 100 percent, Full |
| 62 | Claude Opus 4.6 Anthropic | 66.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'medium' output effort. [variant] claude-opus-4-6-thinking-120K-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 66.3 | 100% confidence 100 percent, Full |
| 63 | GLM-5.3-Flash Z.ai | 65.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] zai-glm-5-3-flash-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.8 | 100% confidence 100 percent, Full |
| 64 | Grok 4.20 (Reasoning) xAI | 65.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] grok-4.20-beta-0309b-reasoningPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.1 | 80% confidence 80 percent, High |
| 65 | Grok 4.6 xAI | 65.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] xai-grok-4-6-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.1 | 100% confidence 100 percent, Full |
| 66 | Claude Opus 4.6 Anthropic | 64.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'low' output effort. [variant] claude-opus-4-6-thinking-120K-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 64.6 | 100% confidence 100 percent, Full |
| 67 | Gemini 3.7 Flash Google | 63.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] google-gemini-3-7-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.7 | 100% confidence 100 percent, Full |
| 68 | Claude Opus 4.8 Anthropic | 62.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' output effort. [variant] anthropic-opus-4-8-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 62.2 | 100% confidence 100 percent, Full |
| 69 | Claude Opus 4.7 Anthropic | 62.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 62.1 | 100% confidence 100 percent, Full |
| 70 | DeepSeek V4 Flash 0731 DeepSeek | 61.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 61.4 | 69% confidence 69 percent, Medium |
| 71 | Grok 4.7 xAI | 61.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 61.4 | 100% confidence 100 percent, Full |
| 72 | DeepSeek V4 Pro 0813 DeepSeek | 61.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-pro-0813-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 61.3 | 69% confidence 69 percent, Medium |
| 73 | Grok 4.6 xAI | 61.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] xai-grok-4-6-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 61.3 | 100% confidence 100 percent, Full |
| 74 | DeepSeek V4.1 Flash DeepSeek | 60.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] deepseek-v4-1-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 60.6 | 69% confidence 69 percent, Medium |
| 75 | Claude Sonnet 4.6 Anthropic | 60.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude_sonnet_4_6_highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 60.4 | 100% confidence 100 percent, Full |
| 76 | Gemini 3.6 Flash Google | 60.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-6-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 60.4 | 100% confidence 100 percent, Full |
| 77 | Kimi K3 Moonshot AI | 60.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 60.4 | 100% confidence 100 percent, Full |
| 78 | DeepSeek V4 Pro 0813 DeepSeek | 59.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-pro-0813-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.7 | 69% confidence 69 percent, Medium |
| 79 | GPT-5.6 Luna OpenAI | 59.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.5 | 100% confidence 100 percent, Full |
| 80 | GPT-6 Luna OpenAI | 59.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.3 | 100% confidence 100 percent, Full |
| 81 | Grok 4.7 xAI | 58.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.8 | 100% confidence 100 percent, Full |
| 82 | Grok 4.7 xAI | 58.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.3 | 100% confidence 100 percent, Full |
| 83 | Claude Sonnet 4.6 Anthropic | 58.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude_sonnet_4_6_maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.3 | 100% confidence 100 percent, Full |
| 84 | GPT-6 Sol OpenAI | 57.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 57.8 | 100% confidence 100 percent, Full |
| 85 | DeepSeek V4 Pro 0813 DeepSeek | 56.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-pro-0813-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 56.3 | 69% confidence 69 percent, Medium |
| 86 | DeepSeek V4 Flash 0731 DeepSeek | 56.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 56.0 | 69% confidence 69 percent, Medium |
| 87 | GPT-5.4 OpenAI | 55.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 55.4 | 100% confidence 100 percent, Full |
| 88 | Kimi K3 Moonshot AI | 55.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 55.0 | 100% confidence 100 percent, Full |
| 89 | GPT-5.2 Pro OpenAI | 54.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 54.2 | 48% confidence 48 percent, Low |
| 90 | Gemini 3.7 Flash Google | 52.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] google-gemini-3-7-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 52.9 | 100% confidence 100 percent, Full |
| 91 | GPT-5.2 OpenAI | 52.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 52.9 | 100% confidence 100 percent, Full |
| 92 | Grok 4.5 xAI | 52.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 52.6 | 100% confidence 100 percent, Full |
| 93 | Grok 4.5 xAI | 52.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 52.6 | 100% confidence 100 percent, Full |
| 94 | Gemini 3.6 Flash Google | 50.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-6-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 50.4 | 100% confidence 100 percent, Full |
| 95 | GLM-5.3-Flash Z.ai | 50.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] zai-glm-5-3-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 50.1 | 100% confidence 100 percent, Full |
| 96 | GPT-5.6 Luna OpenAI | 47.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 47.6 | 100% confidence 100 percent, Full |
| 97 | DeepSeek V4 Flash 0731 DeepSeek | 46.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 46.0 | 69% confidence 69 percent, Medium |
| 98 | GPT-5.2 OpenAI | 43.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 43.3 | 100% confidence 100 percent, Full |
| 99 | GPT-5.6 Sol OpenAI | 42.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 42.5 | 100% confidence 100 percent, Full |
| 100 | Qwen3.8 27B Alibaba / Qwen | 42.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] alibaba-qwen3-8-27b-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 42.4 | 69% confidence 69 percent, Medium |
| 101 | GPT-6 Luna OpenAI | 41.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 41.9 | 100% confidence 100 percent, Full |
| 102 | Inkling Small Thinking Machines Lab | 40.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] thinky-inkling-small-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 40.1 | 100% confidence 100 percent, Full |
| 103 | GPT-5.2 Pro OpenAI | 38.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 38.5 | 48% confidence 48 percent, Low |
| 104 | Claude Opus 4.5 Anthropic | 37.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-64kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 37.6 | 100% confidence 100 percent, Full |
| 105 | GPT-5.6 Terra OpenAI | 37.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 37.5 | 100% confidence 100 percent, Full |
| 106 | Inkling Thinking Machines Lab | 36.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] thinky-inklingPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 36.5 | 100% confidence 100 percent, Full |
| 107 | Gemini 3 Flash Preview Google | 33.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 33.6 | 93% confidence 93 percent, High |
| 108 | GPT-5.5 OpenAI | 33.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 33.3 | 100% confidence 100 percent, Full |
| 109 | Grok 4.5 xAI | 33.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 33.1 | 100% confidence 100 percent, Full |
| 110 | Inkling Small Thinking Machines Lab | 33.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] thinky-inkling-small-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 33.1 | 100% confidence 100 percent, Full |
| 111 | GPT-6 Sol OpenAI | 31.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 31.5 | 100% confidence 100 percent, Full |
| 112 | GPT-6 Luna OpenAI | 31.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 31.4 | 100% confidence 100 percent, Full |
| 113 | Gemini 3.6 Flash Google | 30.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-6-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 30.4 | 100% confidence 100 percent, Full |
| 114 | GPT-5.6 Luna OpenAI | 29.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 29.3 | 100% confidence 100 percent, Full |
| 115 | GPT-5.4 OpenAI | 29.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 29.2 | 100% confidence 100 percent, Full |
| 116 | GLM-5.3-Flash Z.ai | 27.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] zai-glm-5-3-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 27.9 | 100% confidence 100 percent, Full |
| 117 | Grok 4.6 xAI | 27.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] xai-grok-4-6-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 27.6 | 100% confidence 100 percent, Full |
| 118 | GPT-5.2 OpenAI | 26.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 26.7 | 100% confidence 100 percent, Full |
| 119 | Claude Opus 4.5 Anthropic | 22.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 22.8 | 100% confidence 100 percent, Full |
| 120 | Qwen3.8 27B Alibaba / Qwen | 22.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] alibaba-qwen3-8-27b-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 22.8 | 69% confidence 69 percent, Medium |
| 121 | GLM-5.2 Z.ai | 22.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5.2Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 22.8 | 100% confidence 100 percent, Full |
| 122 | GPT-5.4 mini OpenAI | 18.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 18.9 | 100% confidence 100 percent, Full |
| 123 | GPT-5.6 Terra OpenAI | 18.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 18.8 | 100% confidence 100 percent, Full |
| 124 | GPT-5 Pro OpenAI | 18.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-pro-2025-10-06Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 18.3 | 64% confidence 64 percent, Medium |
| 125 | GPT-6 Luna OpenAI | 18.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 18.1 | 100% confidence 100 percent, Full |
| 126 | GPT-5.1 OpenAI | 17.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 17.6 | 85% confidence 85 percent, High |
| 127 | Grok 4.7 xAI | 16.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] xai-grok-4-7-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 16.7 | 100% confidence 100 percent, Full |
| 128 | Claude Opus 4.5 Anthropic | 13.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 13.9 | 100% confidence 100 percent, Full |
| 129 | Inkling Small Thinking Machines Lab | 13.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] thinky-inkling-small-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 13.6 | 100% confidence 100 percent, Full |
| 130 | Claude Sonnet 4.5 (latest) Anthropic | 13.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 13.6 | 43% confidence 43 percent, Low |
| 131 | Qwen3.8 27B Alibaba / Qwen | 13.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] alibaba-qwen3-8-27b-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 13.2 | 69% confidence 69 percent, Medium |
| 132 | GPT-5.4 mini OpenAI | 13.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 13.2 | 100% confidence 100 percent, Full |
| 133 | Gemini 3 Flash Preview Google | 12.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 12.8 | 93% confidence 93 percent, High |
| 134 | Kimi K3 Moonshot AI | 12.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 12.4 | 100% confidence 100 percent, Full |
| 135 | Kimi K2.5 Moonshot AI | 11.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] kimi-k2.5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 11.8 | 100% confidence 100 percent, Full |
| 136 | Gemini 3.5 Flash Lite Google | 10.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-5-flash-lite-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 10.3 | 100% confidence 100 percent, Full |
| 137 | GPT-5 OpenAI | 9.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 9.9 | 100% confidence 100 percent, Full |
| 138 | GPT-5.2 OpenAI | 9.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 9.7 | 100% confidence 100 percent, Full |
| 139 | Claude Opus 4 Anthropic | 8.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 8.6 | 100% confidence 100 percent, Full |
| 140 | Claude Opus 4.5 Anthropic | 7.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.8 | 100% confidence 100 percent, Full |
| 141 | GPT-5 OpenAI | 7.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.5 | 100% confidence 100 percent, Full |
| 142 | GPT-5.6 Luna OpenAI | 7.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.4 | 100% confidence 100 percent, Full |
| 143 | Claude Sonnet 4.5 (latest) Anthropic | 6.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 6.9 | 43% confidence 43 percent, Low |
| 144 | Claude Sonnet 4.5 (latest) Anthropic | 6.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 6.9 | 43% confidence 43 percent, Low |
| 145 | GPT-5.1 OpenAI | 6.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 6.5 | 85% confidence 85 percent, High |
| 146 | o3 OpenAI | 6.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 6.5 | 87% confidence 87 percent, High |
| 147 | o4-mini OpenAI | 6.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 6.1 | 87% confidence 87 percent, High |
| 148 | Claude Sonnet 4 Anthropic | 5.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.9 | 100% confidence 100 percent, Full |
| 149 | Claude Sonnet 4.5 (latest) Anthropic | 5.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.8 | 43% confidence 43 percent, Low |
| 150 | GPT-5.4 nano OpenAI | 5.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.7 | 100% confidence 100 percent, Full |
| 151 | Gemini 3.5 Flash Lite Google | 5.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-5-flash-lite-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.3 | 100% confidence 100 percent, Full |
| 152 | GPT-5.6 Luna OpenAI | 5.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.1 | 100% confidence 100 percent, Full |
| 153 | Inkling Small Thinking Machines Lab | 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] thinky-inkling-small-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.9 | 100% confidence 100 percent, Full |
| 154 | Gemini 2.5 Pro Google | 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.9 | 90% confidence 90 percent, High |
| 155 | MiniMax-M2.5 MiniMax | 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] minimax-m2.5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.9 | 100% confidence 100 percent, Full |
| 156 | o3-pro OpenAI | 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.9 | 22% confidence 22 percent, Low |
| 157 | GLM-5 Z.ai | 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.9 | 90% confidence 90 percent, High |
| 158 | GPT-6 Luna OpenAI | 4.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.6 | 100% confidence 100 percent, Full |
| 159 | Claude Opus 4 Anthropic | 4.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.5 | 100% confidence 100 percent, Full |
| 160 | GPT-5 Mini OpenAI | 4.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.4 | 99% confidence 99 percent, High |
| 161 | GPT-5.4 mini OpenAI | 4.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.4 | 100% confidence 100 percent, Full |
| 162 | Claude Haiku 4.5 (latest) Anthropic | 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.0 | 33% confidence 33 percent, Low |
| 163 | DeepSeek V3.2 DeepSeek | 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek-v3.2Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.0 | 56% confidence 56 percent, Medium |
| 164 | Gemini 2.5 Pro Google | 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.0 | 90% confidence 90 percent, High |
| 165 | GPT-5 Mini OpenAI | 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.0 | 99% confidence 99 percent, High |
| 166 | Claude Sonnet 4.5 (latest) Anthropic | 3.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.8 | 43% confidence 43 percent, Low |
| 167 | GPT-5.4 nano OpenAI | 3.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.6 | 100% confidence 100 percent, Full |
| 168 | o3-mini OpenAI | 3.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.0 | 96% confidence 96 percent, High |
| 169 | o3 OpenAI | 3.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.0 | 87% confidence 87 percent, High |
| 170 | Gemini 2.5 Pro Google | 2.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.9 | 90% confidence 90 percent, High |
| 171 | Claude Haiku 4.5 (latest) Anthropic | 2.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.8 | 33% confidence 33 percent, Low |
| 172 | GPT-5 Nano OpenAI | 2.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.6 | 85% confidence 85 percent, High |
| 173 | o4-mini OpenAI | 2.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.4 | 87% confidence 87 percent, High |
| 174 | Claude Sonnet 4 Anthropic | 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.1 | 100% confidence 100 percent, Full |
| 175 | o3-mini OpenAI | 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.1 | 96% confidence 96 percent, High |
| 176 | o3-pro OpenAI | 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.1 | 22% confidence 22 percent, Low |
| 177 | o3 OpenAI | 2.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.0 | 87% confidence 87 percent, High |
| 178 | GPT-5 OpenAI | 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.9 | 100% confidence 100 percent, Full |
| 179 | GPT-5.1 OpenAI | 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.9 | 85% confidence 85 percent, High |
| 180 | GPT-5.4 nano OpenAI | 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.9 | 100% confidence 100 percent, Full |
| 181 | o3-pro OpenAI | 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.9 | 22% confidence 22 percent, Low |
| 182 | Claude Haiku 4.5 (latest) Anthropic | 1.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.7 | 33% confidence 33 percent, Low |
| 183 | o4-mini OpenAI | 1.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.7 | 87% confidence 87 percent, High |
| 184 | GPT-5.4 nano OpenAI | 1.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.5 | 100% confidence 100 percent, Full |
| 185 | Gemini 3.5 Flash Lite Google | 1.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-5-flash-lite-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.5 | 100% confidence 100 percent, Full |
| 186 | DeepSeek-R1 DeepSeek | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] R1Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 88% confidence 88 percent, High |
| 187 | Gemini 2.0 Flash Google | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Gemini 2.0 FlashPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 28% confidence 28 percent, Low |
| 188 | Claude Opus 4 Anthropic | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 100% confidence 100 percent, Full |
| 189 | Claude Sonnet 4 Anthropic | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 100% confidence 100 percent, Full |
| 190 | Qwen3 235B-A22B Instruct 2507 Alibaba / Qwen | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] qwen3-235b-a22b-instruct-2507Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 48% confidence 48 percent, Low |
| 191 | Claude Haiku 4.5 (latest) Anthropic | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 33% confidence 33 percent, Low |
| 192 | Claude Haiku 4.5 (latest) Anthropic | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 33% confidence 33 percent, Low |
| 193 | Gemini 3 Flash Preview Google | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 93% confidence 93 percent, High |
| 194 | DeepSeek-R1 DeepSeek | 1.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.1 | 88% confidence 88 percent, High |
| 195 | GPT-5.4 mini OpenAI | 1.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.1 | 100% confidence 100 percent, Full |
| 196 | Claude Sonnet 3.7 Anthropic | 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 8KPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.9 | 100% confidence 100 percent, Full |
| 197 | GPT-5 Nano OpenAI | 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.9 | 85% confidence 85 percent, High |
| 198 | Claude Sonnet 4 Anthropic | 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.9 | 100% confidence 100 percent, Full |
| 199 | GPT-5 Mini OpenAI | 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.8 | 99% confidence 99 percent, High |
| 200 | GPT-5.2 OpenAI | 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.8 | 100% confidence 100 percent, Full |
| 201 | Claude Sonnet 3.7 Anthropic | 0.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.7 | 100% confidence 100 percent, Full |
| 202 | GPT-4.1 OpenAI | 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.4 | 100% confidence 100 percent, Full |
| 203 | GPT-5.1 OpenAI | 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.4 | 85% confidence 85 percent, High |
| 204 | Claude Sonnet 3.7 Anthropic | 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 1KPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.4 | 100% confidence 100 percent, Full |
| 205 | Claude Sonnet 3.7 Anthropic | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 206 | Claude Opus 4 Anthropic | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 207 | Gemini 2.5 Pro Google | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 90% confidence 90 percent, High |
| 208 | Llama 4 Maverick 17B Instruct Meta | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Llama-4-Maverick-17B-128E-Instruct-FP8-togetherPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 78% confidence 78 percent, Medium |
| 209 | Magistral Medium (latest) Mistral AI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 16% confidence 16 percent, Low |
| 210 | Magistral Medium (latest) Mistral AI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506-thinkingPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 16% confidence 16 percent, Low |
| 211 | Magistral Small Mistral AI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 48% confidence 48 percent, Low |
| 212 | GPT-4.1 mini OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-mini-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 87% confidence 87 percent, High |
| 213 | GPT-4.1 nano OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-nano-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 82% confidence 82 percent, High |
| 214 | GPT-4o OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4o-2024-11-20Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 59% confidence 59 percent, Medium |
| 215 | GPT-4o mini OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4o-mini-2024-07-18Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 216 | GPT-5 Nano OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 85% confidence 85 percent, High |
| 217 | o3-mini OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 96% confidence 96 percent, High |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.