ARC-AGI-2 (public eval)Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Claude Fable 5.1 Anthropic | 99.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-1-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 99.2 | 100% confidence 100 percent, Full |
| 2 | Claude Fable 5.1 Anthropic | 98.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-1-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 98.3 | 100% confidence 100 percent, Full |
| 3 | Claude Opus 5.5 Anthropic | 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.9 | 100% confidence 100 percent, Full |
| 4 | GPT-6 Astra OpenAI | 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.9 | 100% confidence 100 percent, Full |
| 5 | GPT-6 Astra OpenAI | 97.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.9 | 100% confidence 100 percent, Full |
| 6 | Claude Opus 5.5 Anthropic | 97.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.6 | 100% confidence 100 percent, Full |
| 7 | GPT-6 Astra OpenAI | 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.5 | 100% confidence 100 percent, Full |
| 8 | GPT-6.1 Sol OpenAI | 97.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.1 | 100% confidence 100 percent, Full |
| 9 | Claude Opus 5 Anthropic | 97.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 97.1 | 100% confidence 100 percent, Full |
| 10 | GPT-6 Astra OpenAI | 96.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.7 | 100% confidence 100 percent, Full |
| 11 | Claude Fable 5 Anthropic | 96.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 96.4 | 100% confidence 100 percent, Full |
| 12 | GPT-6.1 Sol OpenAI | 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.8 | 100% confidence 100 percent, Full |
| 13 | GPT-6.1 Sol OpenAI | 95.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 95.8 | 100% confidence 100 percent, Full |
| 14 | Claude Fable 5.1 Anthropic | 94.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-1-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.7 | 100% confidence 100 percent, Full |
| 15 | GPT-6 Astra OpenAI | 94.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 94.2 | 100% confidence 100 percent, Full |
| 16 | Claude Fable 5 Anthropic | 93.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.9 | 100% confidence 100 percent, Full |
| 17 | GPT-5.6 Sol OpenAI | 93.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.8 | 100% confidence 100 percent, Full |
| 18 | Claude Opus 5.5 Anthropic | 93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.5 | 100% confidence 100 percent, Full |
| 19 | Claude Opus 5.5 Anthropic | 93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.5 | 100% confidence 100 percent, Full |
| 20 | Claude Fable 5 Anthropic | 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.3 | 100% confidence 100 percent, Full |
| 21 | Claude Opus 5 Anthropic | 93.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 93.2 | 100% confidence 100 percent, Full |
| 22 | Gemini 3.8 Flash Google | 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.5 | 100% confidence 100 percent, Full |
| 23 | GPT-5.4 Pro OpenAI | 92.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-pro-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 92.2 | 69% confidence 69 percent, Medium |
| 24 | GPT-5.6 Sol OpenAI | 91.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.9 | 100% confidence 100 percent, Full |
| 25 | GPT-5.6 Terra OpenAI | 91.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 91.3 | 100% confidence 100 percent, Full |
| 26 | GPT-5.5 OpenAI | 90.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.6 | 100% confidence 100 percent, Full |
| 27 | GPT-6.1 Sol OpenAI | 90.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.6 | 100% confidence 100 percent, Full |
| 28 | GPT-5.5 Pro OpenAI | 90.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.4 | 53% confidence 53 percent, Medium |
| 29 | GPT-5.5 Pro OpenAI | 90.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.1 | 53% confidence 53 percent, Medium |
| 30 | GPT-6 Sol OpenAI | 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 90.0 | 100% confidence 100 percent, Full |
| 31 | Gemini 3.1 Pro Preview Google | 88.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-1-pro-previewPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 88.1 | 100% confidence 100 percent, Full |
| 32 | Claude Fable 5 Anthropic | 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.5 | 100% confidence 100 percent, Full |
| 33 | Claude Fable 5.1 Anthropic | 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-1-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.5 | 100% confidence 100 percent, Full |
| 34 | Gemini 3.8 Flash Google | 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 87.5 | 100% confidence 100 percent, Full |
| 35 | GPT-5.6 Sol OpenAI | 86.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.1 | 100% confidence 100 percent, Full |
| 36 | Claude Fable 5.1 Anthropic | 86.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-1-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 86.0 | 100% confidence 100 percent, Full |
| 37 | Gemini 3.7 Flash Google | 85.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] google-gemini-3-7-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 85.3 | 100% confidence 100 percent, Full |
| 38 | GPT-5.4 OpenAI | 84.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 84.2 | 100% confidence 100 percent, Full |
| 39 | GPT-5.6 Terra OpenAI | 82.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 82.6 | 100% confidence 100 percent, Full |
| 40 | Claude Opus 4.7 Anthropic | 82.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 82.2 | 100% confidence 100 percent, Full |
| 41 | GPT-5.5 OpenAI | 82.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 82.2 | 100% confidence 100 percent, Full |
| 42 | Claude Opus 4.7 Anthropic | 81.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 81.9 | 100% confidence 100 percent, Full |
| 43 | GPT-6 Sol OpenAI | 81.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 81.8 | 100% confidence 100 percent, Full |
| 44 | DeepSeek V4.1 Flash DeepSeek | 81.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 81.7 | 69% confidence 69 percent, Medium |
| 45 | Claude Opus 4.7 Anthropic | 80.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 80.7 | 100% confidence 100 percent, Full |
| 46 | Claude Fable 5 Anthropic | 80.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 80.4 | 100% confidence 100 percent, Full |
| 47 | Claude Opus 4.6 Anthropic | 79.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude-opus-4-6-thinking-120K-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 79.0 | 100% confidence 100 percent, Full |
| 48 | GPT-6.1 Sol OpenAI | 77.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 77.8 | 100% confidence 100 percent, Full |
| 49 | Gemini 3.8 Flash Google | 77.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 77.2 | 100% confidence 100 percent, Full |
| 50 | GPT-5.4 OpenAI | 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 75.8 | 100% confidence 100 percent, Full |
| 51 | Claude Opus 4.6 Anthropic | 74.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude-opus-4-6-thinking-120K-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 74.9 | 100% confidence 100 percent, Full |
| 52 | GPT-5.6 Terra OpenAI | 73.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 73.9 | 100% confidence 100 percent, Full |
| 53 | Claude Opus 4.6 Anthropic | 73.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'medium' output effort. [variant] claude-opus-4-6-thinking-120K-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 73.6 | 100% confidence 100 percent, Full |
| 54 | Gemini 3.5 Flash Google | 72.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 72.1 | 100% confidence 100 percent, Full |
| 55 | DeepSeek V4.1 Flash DeepSeek | 71.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 71.9 | 69% confidence 69 percent, Medium |
| 56 | Claude Opus 4.7 Anthropic | 71.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 71.6 | 100% confidence 100 percent, Full |
| 57 | GPT-5.5 OpenAI | 70.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 70.7 | 100% confidence 100 percent, Full |
| 58 | GPT-5.6 Sol OpenAI | 70.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 70.7 | 100% confidence 100 percent, Full |
| 59 | Gemini 3.7 Flash Google | 68.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] google-gemini-3-7-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.6 | 100% confidence 100 percent, Full |
| 60 | Grok 4.6 xAI | 68.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] xai-grok-4-6-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.3 | 100% confidence 100 percent, Full |
| 61 | Grok 4.6 xAI | 68.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] xai-grok-4-6-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 68.3 | 100% confidence 100 percent, Full |
| 62 | Claude Opus 5.5 Anthropic | 67.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.8 | 100% confidence 100 percent, Full |
| 63 | Grok 4.6 xAI | 67.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] xai-grok-4-6-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 67.4 | 100% confidence 100 percent, Full |
| 64 | DeepSeek V4.1 Flash DeepSeek | 66.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] deepseek-v4-1-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 66.8 | 69% confidence 69 percent, Medium |
| 65 | GLM-5.3-Flash Z.ai | 66.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] zai-glm-5-3-flash-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 66.8 | 100% confidence 100 percent, Full |
| 66 | Claude Sonnet 4.6 Anthropic | 65.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude_sonnet_4_6_highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.7 | 100% confidence 100 percent, Full |
| 67 | Gemini 3.6 Flash Google | 65.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-6-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 65.3 | 100% confidence 100 percent, Full |
| 68 | DeepSeek V4 Flash 0731 DeepSeek | 63.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.9 | 69% confidence 69 percent, Medium |
| 69 | GPT-6 Sol OpenAI | 63.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.6 | 100% confidence 100 percent, Full |
| 70 | DeepSeek V4 Pro 0813 DeepSeek | 63.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-pro-0813-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.6 | 69% confidence 69 percent, Medium |
| 71 | Grok 4.20 (Reasoning) xAI | 63.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] grok-4.20-beta-0309b-reasoningPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 63.6 | 80% confidence 80 percent, High |
| 72 | Claude Sonnet 4.6 Anthropic | 62.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude_sonnet_4_6_maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 62.4 | 100% confidence 100 percent, Full |
| 73 | GPT-6 Luna OpenAI | 61.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 61.8 | 100% confidence 100 percent, Full |
| 74 | Kimi K3 Moonshot AI | 61.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 61.7 | 100% confidence 100 percent, Full |
| 75 | GPT-5.6 Luna OpenAI | 60.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 60.2 | 100% confidence 100 percent, Full |
| 76 | Grok 4.7 xAI | 60%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 60.0 | 100% confidence 100 percent, Full |
| 77 | DeepSeek V4 Flash 0731 DeepSeek | 59.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.9 | 69% confidence 69 percent, Medium |
| 78 | Claude Opus 4.6 Anthropic | 59.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'low' output effort. [variant] claude-opus-4-6-thinking-120K-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.9 | 100% confidence 100 percent, Full |
| 79 | GPT-5.2 OpenAI | 59.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.8 | 100% confidence 100 percent, Full |
| 80 | DeepSeek V4 Pro 0813 DeepSeek | 59.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-pro-0813-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.4 | 69% confidence 69 percent, Medium |
| 81 | DeepSeek V4 Pro 0813 DeepSeek | 58.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-pro-0813-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.9 | 69% confidence 69 percent, Medium |
| 82 | GPT-5.4 OpenAI | 58.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.2 | 100% confidence 100 percent, Full |
| 83 | Grok 4.5 xAI | 58.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 58.2 | 100% confidence 100 percent, Full |
| 84 | Grok 4.7 xAI | 57.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 57.9 | 100% confidence 100 percent, Full |
| 85 | Grok 4.5 xAI | 57.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 57.5 | 100% confidence 100 percent, Full |
| 86 | Grok 4.7 xAI | 57.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 57.5 | 100% confidence 100 percent, Full |
| 87 | Gemini 3.6 Flash Google | 56.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-6-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 56.5 | 100% confidence 100 percent, Full |
| 88 | Gemini 3.7 Flash Google | 54.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] google-gemini-3-7-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 54.7 | 100% confidence 100 percent, Full |
| 89 | Kimi K3 Moonshot AI | 53.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 53.2 | 100% confidence 100 percent, Full |
| 90 | GLM-5.3-Flash Z.ai | 51.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] zai-glm-5-3-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 51.9 | 100% confidence 100 percent, Full |
| 91 | GPT-5.2 Pro OpenAI | 51.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 51.7 | 48% confidence 48 percent, Low |
| 92 | GPT-5.6 Luna OpenAI | 51.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 51.4 | 100% confidence 100 percent, Full |
| 93 | GPT-6 Sol OpenAI | 49.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 49.4 | 100% confidence 100 percent, Full |
| 94 | DeepSeek V4 Flash 0731 DeepSeek | 47.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 47.4 | 69% confidence 69 percent, Medium |
| 95 | GPT-5.2 OpenAI | 39.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 39.9 | 100% confidence 100 percent, Full |
| 96 | Inkling Small Thinking Machines Lab | 39.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] thinky-inkling-small-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 39.3 | 100% confidence 100 percent, Full |
| 97 | GPT-5.6 Terra OpenAI | 38.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 38.9 | 100% confidence 100 percent, Full |
| 98 | Qwen3.8 27B Alibaba / Qwen | 38.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] alibaba-qwen3-8-27b-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 38.8 | 69% confidence 69 percent, Medium |
| 99 | GPT-5.6 Sol OpenAI | 38.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 38.5 | 100% confidence 100 percent, Full |
| 100 | GPT-5.2 Pro OpenAI | 37.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 37.9 | 48% confidence 48 percent, Low |
| 101 | Inkling Thinking Machines Lab | 37.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] thinky-inklingPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 37.8 | 100% confidence 100 percent, Full |
| 102 | Grok 4.5 xAI | 36.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 36.3 | 100% confidence 100 percent, Full |
| 103 | Grok 4.6 xAI | 36.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] xai-grok-4-6-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 36.3 | 100% confidence 100 percent, Full |
| 104 | GPT-6 Luna OpenAI | 35.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 35.1 | 100% confidence 100 percent, Full |
| 105 | GPT-5.5 OpenAI | 34.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 34.3 | 100% confidence 100 percent, Full |
| 106 | Gemini 3 Flash Preview Google | 34.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 34.0 | 93% confidence 93 percent, High |
| 107 | GPT-6 Luna OpenAI | 31.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 31.9 | 100% confidence 100 percent, Full |
| 108 | Gemini 3.6 Flash Google | 31.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-6-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 31.1 | 100% confidence 100 percent, Full |
| 109 | GPT-5.6 Luna OpenAI | 30.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 30.1 | 100% confidence 100 percent, Full |
| 110 | Claude Opus 4.5 Anthropic | 28.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 28.1 | 100% confidence 100 percent, Full |
| 111 | GPT-5.2 OpenAI | 27.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 27.6 | 100% confidence 100 percent, Full |
| 112 | Inkling Small Thinking Machines Lab | 25.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] thinky-inkling-small-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 25.7 | 100% confidence 100 percent, Full |
| 113 | GLM-5.3-Flash Z.ai | 25.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] zai-glm-5-3-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 25.7 | 100% confidence 100 percent, Full |
| 114 | GPT-6 Sol OpenAI | 25.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 25.6 | 100% confidence 100 percent, Full |
| 115 | Grok 4.7 xAI | 25%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] xai-grok-4-7-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 25.0 | 100% confidence 100 percent, Full |
| 116 | Claude Opus 4.5 Anthropic | 24.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 24.2 | 100% confidence 100 percent, Full |
| 117 | GPT-5.4 OpenAI | 23.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 23.2 | 100% confidence 100 percent, Full |
| 118 | GLM-5.2 Z.ai | 20.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5.2Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 20.8 | 100% confidence 100 percent, Full |
| 119 | Qwen3.8 27B Alibaba / Qwen | 19.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] alibaba-qwen3-8-27b-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 19.6 | 69% confidence 69 percent, Medium |
| 120 | GPT-6 Luna OpenAI | 19.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 19.2 | 100% confidence 100 percent, Full |
| 121 | GPT-5.1 OpenAI | 18.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 18.3 | 85% confidence 85 percent, High |
| 122 | GPT-5.4 mini OpenAI | 17.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 17.8 | 100% confidence 100 percent, Full |
| 123 | Qwen3.8 27B Alibaba / Qwen | 17.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] alibaba-qwen3-8-27b-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 17.1 | 69% confidence 69 percent, Medium |
| 124 | Kimi K3 Moonshot AI | 16.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 16.1 | 100% confidence 100 percent, Full |
| 125 | Gemini 3 Flash Preview Google | 15.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 15.3 | 93% confidence 93 percent, High |
| 126 | GPT-5.6 Terra OpenAI | 14.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 14.9 | 100% confidence 100 percent, Full |
| 127 | Claude Sonnet 4.5 (latest) Anthropic | 14.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 14.7 | 43% confidence 43 percent, Low |
| 128 | Inkling Small Thinking Machines Lab | 13.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] thinky-inkling-small-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 13.6 | 100% confidence 100 percent, Full |
| 129 | GPT-5 Pro OpenAI | 13.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-pro-2025-10-06Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 13.3 | 64% confidence 64 percent, Medium |
| 130 | Kimi K2.5 Moonshot AI | 12.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] kimi-k2.5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 12.1 | 100% confidence 100 percent, Full |
| 131 | Claude Opus 4.5 Anthropic | 10.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 10.4 | 100% confidence 100 percent, Full |
| 132 | GPT-5 OpenAI | 9.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 9.6 | 100% confidence 100 percent, Full |
| 133 | GPT-5.1 OpenAI | 8.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 8.4 | 85% confidence 85 percent, High |
| 134 | GPT-5.2 OpenAI | 8.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 8.3 | 100% confidence 100 percent, Full |
| 135 | Gemini 3.5 Flash Lite Google | 7.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-5-flash-lite-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.9 | 100% confidence 100 percent, Full |
| 136 | GPT-5.6 Luna OpenAI | 7.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.6 | 100% confidence 100 percent, Full |
| 137 | GPT-5 OpenAI | 7.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.6 | 100% confidence 100 percent, Full |
| 138 | o4-mini OpenAI | 7.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.5 | 87% confidence 87 percent, High |
| 139 | Claude Opus 4.5 Anthropic | 7.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.1 | 100% confidence 100 percent, Full |
| 140 | GPT-5.4 mini OpenAI | 7.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.0 | 100% confidence 100 percent, Full |
| 141 | Claude Sonnet 4.5 (latest) Anthropic | 6.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 6.5 | 43% confidence 43 percent, Low |
| 142 | Inkling Small Thinking Machines Lab | 6.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] thinky-inkling-small-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 6.4 | 100% confidence 100 percent, Full |
| 143 | GPT-5 Mini OpenAI | 5.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.8 | 99% confidence 99 percent, High |
| 144 | MiniMax-M2.5 MiniMax | 5.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] minimax-m2.5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.4 | 100% confidence 100 percent, Full |
| 145 | GPT-5.4 mini OpenAI | 5.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.4 | 100% confidence 100 percent, Full |
| 146 | GLM-5 Z.ai | 5.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.4 | 90% confidence 90 percent, High |
| 147 | Claude Haiku 4.5 (latest) Anthropic | 5.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.1 | 33% confidence 33 percent, Low |
| 148 | Gemini 2.5 Pro Google | 5.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.1 | 90% confidence 90 percent, High |
| 149 | GPT-5.4 nano OpenAI | 5.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 5.1 | 100% confidence 100 percent, Full |
| 150 | Claude Sonnet 4.5 (latest) Anthropic | 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.9 | 43% confidence 43 percent, Low |
| 151 | Gemini 2.5 Pro Google | 4.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-32kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.6 | 90% confidence 90 percent, High |
| 152 | o3 OpenAI | 4.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.5 | 87% confidence 87 percent, High |
| 153 | Claude Sonnet 4 Anthropic | 4.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.4 | 100% confidence 100 percent, Full |
| 154 | Claude Haiku 4.5 (latest) Anthropic | 4.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.2 | 33% confidence 33 percent, Low |
| 155 | GPT-5.4 nano OpenAI | 4.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.2 | 100% confidence 100 percent, Full |
| 156 | DeepSeek V3.2 DeepSeek | 3.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek-v3.2Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.9 | 56% confidence 56 percent, Medium |
| 157 | o3-pro OpenAI | 3.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.9 | 22% confidence 22 percent, Low |
| 158 | Claude Sonnet 4.5 (latest) Anthropic | 3.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.8 | 43% confidence 43 percent, Low |
| 159 | GPT-5.4 nano OpenAI | 3.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.5 | 100% confidence 100 percent, Full |
| 160 | o3-pro OpenAI | 3.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.5 | 22% confidence 22 percent, Low |
| 161 | Claude Opus 4 Anthropic | 3.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-16kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.3 | 100% confidence 100 percent, Full |
| 162 | Claude Sonnet 4 Anthropic | 2.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.9 | 100% confidence 100 percent, Full |
| 163 | o3 OpenAI | 2.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.9 | 87% confidence 87 percent, High |
| 164 | GPT-5.6 Luna OpenAI | 2.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.8 | 100% confidence 100 percent, Full |
| 165 | o3 OpenAI | 2.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.7 | 87% confidence 87 percent, High |
| 166 | Gemini 2.5 Pro Google | 2.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.6 | 90% confidence 90 percent, High |
| 167 | Claude Opus 4 Anthropic | 2.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.5 | 100% confidence 100 percent, Full |
| 168 | GPT-5 OpenAI | 2.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.5 | 100% confidence 100 percent, Full |
| 169 | GPT-5.1 OpenAI | 2.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.2 | 85% confidence 85 percent, High |
| 170 | o4-mini OpenAI | 2.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.2 | 87% confidence 87 percent, High |
| 171 | Claude Haiku 4.5 (latest) Anthropic | 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-8kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.1 | 33% confidence 33 percent, Low |
| 172 | Claude Sonnet 4 Anthropic | 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.1 | 100% confidence 100 percent, Full |
| 173 | Claude Sonnet 4.5 (latest) Anthropic | 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.1 | 43% confidence 43 percent, Low |
| 174 | Gemini 3.5 Flash Lite Google | 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-5-flash-lite-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.9 | 100% confidence 100 percent, Full |
| 175 | o3-pro OpenAI | 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.9 | 22% confidence 22 percent, Low |
| 176 | Claude Opus 4 Anthropic | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 100% confidence 100 percent, Full |
| 177 | Claude Sonnet 4 Anthropic | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 100% confidence 100 percent, Full |
| 178 | Gemini 3 Flash Preview Google | 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.3 | 93% confidence 93 percent, High |
| 179 | GPT-6 Luna OpenAI | 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.8 | 100% confidence 100 percent, Full |
| 180 | Qwen3 235B-A22B Instruct 2507 Alibaba / Qwen | 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] qwen3-235b-a22b-instruct-2507Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.8 | 48% confidence 48 percent, Low |
| 181 | GPT-5.4 mini OpenAI | 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.8 | 100% confidence 100 percent, Full |
| 182 | Claude Sonnet 3.7 Anthropic | 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.8 | 100% confidence 100 percent, Full |
| 183 | Claude Sonnet 3.7 Anthropic | 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 1KPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.8 | 100% confidence 100 percent, Full |
| 184 | GPT-5 Mini OpenAI | 0.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.6 | 99% confidence 99 percent, High |
| 185 | Claude Opus 4 Anthropic | 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.4 | 100% confidence 100 percent, Full |
| 186 | Gemini 2.5 Pro Google | 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.4 | 90% confidence 90 percent, High |
| 187 | DeepSeek-R1 DeepSeek | 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] R1Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.3 | 88% confidence 88 percent, High |
| 188 | GPT-5 Nano OpenAI | 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.3 | 85% confidence 85 percent, High |
| 189 | DeepSeek-R1 DeepSeek | 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.3 | 88% confidence 88 percent, High |
| 190 | GPT-5 Mini OpenAI | 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.3 | 99% confidence 99 percent, High |
| 191 | o4-mini OpenAI | 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.3 | 87% confidence 87 percent, High |
| 192 | Claude Sonnet 3.7 Anthropic | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 193 | Claude Sonnet 3.7 Anthropic | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 8KPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 194 | Claude Haiku 4.5 (latest) Anthropic | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 33% confidence 33 percent, Low |
| 195 | Claude Haiku 4.5 (latest) Anthropic | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 33% confidence 33 percent, Low |
| 196 | Gemini 3.5 Flash Lite Google | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-5-flash-lite-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 197 | Llama 4 Maverick 17B Instruct Meta | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Llama-4-Maverick-17B-128E-Instruct-FP8-togetherPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 78% confidence 78 percent, Medium |
| 198 | Magistral Medium (latest) Mistral AI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 16% confidence 16 percent, Low |
| 199 | Magistral Medium (latest) Mistral AI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506-thinkingPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 16% confidence 16 percent, Low |
| 200 | Magistral Small Mistral AI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 48% confidence 48 percent, Low |
| 201 | GPT-4.1 OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 202 | GPT-4.1 mini OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-mini-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 87% confidence 87 percent, High |
| 203 | GPT-4.1 nano OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-nano-2025-04-14Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 82% confidence 82 percent, High |
| 204 | GPT-4o OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4o-2024-11-20Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 59% confidence 59 percent, Medium |
| 205 | GPT-4o mini OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4o-mini-2024-07-18Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 206 | GPT-5 Nano OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 85% confidence 85 percent, High |
| 207 | GPT-5 Nano OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 85% confidence 85 percent, High |
| 208 | GPT-5.1 OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 85% confidence 85 percent, High |
| 209 | GPT-5.2 OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-nonePublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 210 | GPT-5.4 nano OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 211 | o3-mini OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 96% confidence 96 percent, High |
| 212 | o3-mini OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 96% confidence 96 percent, High |
| 213 | o3-mini OpenAI | 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 96% confidence 96 percent, High |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.