ARC-AGI-1 (semi-private)Benchmark scores and sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 Claude Fable 5 Anthropic 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
2 Claude Fable 5 Anthropic 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
3 Claude Opus 5.5 Anthropic 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
4 Gemini 3.8 Flash Google 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
5 GPT-6 Astra OpenAI 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
6 GPT-6 Astra OpenAI 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
7 GPT-6.1 Sol OpenAI 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
8 GPT-6.1 Sol OpenAI 98.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.5 100% confidence 100 percent, Full
9 Gemini 3.1 Pro Preview Google 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-1-pro-previewPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.0 100% confidence 100 percent, Full
10 Claude Fable 5.1 Anthropic 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-1-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
11 Claude Opus 5 Anthropic 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
12 Claude Opus 5 Anthropic 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
13 Claude Opus 5.5 Anthropic 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
14 Claude Opus 5.5 Anthropic 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
15 Claude Opus 5.5 Anthropic 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
16 Gemini 3.8 Flash Google 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
17 GPT-5.6 Sol OpenAI 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
18 GPT-6 Astra OpenAI 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
19 GPT-6 Astra OpenAI 97.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.5 100% confidence 100 percent, Full
20 GPT-5.6 Sol OpenAI 97%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
97.0 100% confidence 100 percent, Full
21 Claude Fable 5.1 Anthropic 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-1-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 100% confidence 100 percent, Full
22 GPT-5.5 Pro OpenAI 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 53% confidence 53 percent, Medium
23 GPT-5.6 Sol OpenAI 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 100% confidence 100 percent, Full
24 GPT-5.6 Terra OpenAI 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 100% confidence 100 percent, Full
25 GPT-6 Astra OpenAI 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 100% confidence 100 percent, Full
26 GPT-6.1 Sol OpenAI 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 100% confidence 100 percent, Full
27 Claude Fable 5.1 Anthropic 96%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-1-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.0 100% confidence 100 percent, Full
28 Claude Fable 5 Anthropic 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.5 100% confidence 100 percent, Full
29 Gemini 3.7 Flash Google 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] google-gemini-3-7-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.5 100% confidence 100 percent, Full
30 GPT-6 Sol OpenAI 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.5 100% confidence 100 percent, Full
31 GPT-6.1 Sol OpenAI 95.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.5 100% confidence 100 percent, Full
32 GPT-5.5 OpenAI 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.0 100% confidence 100 percent, Full
33 GPT-5.5 Pro OpenAI 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.0 53% confidence 53 percent, Medium
34 Claude Fable 5.1 Anthropic 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-1-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.5 100% confidence 100 percent, Full
35 DeepSeek V4.1 Flash DeepSeek 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.5 69% confidence 69 percent, Medium
36 Kimi K3 Moonshot AI 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.5 100% confidence 100 percent, Full
37 GPT-5.4 Pro OpenAI 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-pro-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.5 69% confidence 69 percent, Medium
38 GPT-5.5 OpenAI 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.5 100% confidence 100 percent, Full
39 Claude Opus 4.6 Anthropic 94%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude-opus-4-6-thinking-120K-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.0 100% confidence 100 percent, Full
40 GPT-5.6 Terra OpenAI 94%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.0 100% confidence 100 percent, Full
41 GPT-5.4 OpenAI 93.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
93.7 100% confidence 100 percent, Full
42 Claude Opus 4.7 Anthropic 93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
93.5 100% confidence 100 percent, Full
43 GPT-6.1 Sol OpenAI 93.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
93.5 100% confidence 100 percent, Full
44 Claude Opus 4.6 Anthropic 93%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude-opus-4-6-thinking-120K-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
93.0 100% confidence 100 percent, Full
45 GPT-5.4 OpenAI 92.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.7 100% confidence 100 percent, Full
46 GPT-6 Sol OpenAI 92.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.7 100% confidence 100 percent, Full
47 Claude Fable 5 Anthropic 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 100% confidence 100 percent, Full
48 Claude Opus 4.8 Anthropic 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-opus-4-8-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 100% confidence 100 percent, Full
49 Gemini 3.5 Flash Google 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 100% confidence 100 percent, Full
50 GPT-5.6 Sol OpenAI 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 100% confidence 100 percent, Full
51 GPT-5.5 OpenAI 92.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.2 100% confidence 100 percent, Full
52 Claude Opus 4.6 Anthropic 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'medium' output effort. [variant] claude-opus-4-6-thinking-120K-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.0 100% confidence 100 percent, Full
53 Claude Opus 4.7 Anthropic 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.0 100% confidence 100 percent, Full
54 Claude Opus 4.8 Anthropic 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-opus-4-8-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.0 100% confidence 100 percent, Full
55 GPT-5.6 Terra OpenAI 92%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.0 100% confidence 100 percent, Full
56 Claude Opus 4.8 Anthropic 91.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' output effort. [variant] anthropic-opus-4-8-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.5 100% confidence 100 percent, Full
57 Gemini 3.6 Flash Google 91.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-6-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.2 100% confidence 100 percent, Full
58 Gemini 3.7 Flash Google 91.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] google-gemini-3-7-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.2 100% confidence 100 percent, Full
59 Claude Opus 4.7 Anthropic 91%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.0 100% confidence 100 percent, Full
60 Claude Opus 4.7 Anthropic 91%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.0 100% confidence 100 percent, Full
61 GPT-6 Sol OpenAI 91%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.0 100% confidence 100 percent, Full
62 GLM-5.3-Flash Z.ai 91%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] zai-glm-5-3-flash-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.0 100% confidence 100 percent, Full
63 Claude Fable 5 Anthropic 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.5 100% confidence 100 percent, Full
64 DeepSeek V4 Pro 0813 DeepSeek 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-pro-0813-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.5 69% confidence 69 percent, Medium
65 DeepSeek V4.1 Flash DeepSeek 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] deepseek-v4-1-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.5 69% confidence 69 percent, Medium
66 Gemini 3.8 Flash Google 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.5 100% confidence 100 percent, Full
67 GPT-5.2 Pro OpenAI 90.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.5 48% confidence 48 percent, Low
68 Grok 4.7 xAI 90.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.2 100% confidence 100 percent, Full
69 Claude Fable 5.1 Anthropic 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-1-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.0 100% confidence 100 percent, Full
70 DeepSeek V4 Pro 0813 DeepSeek 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-pro-0813-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.0 69% confidence 69 percent, Medium
71 Grok 4.7 xAI 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.0 100% confidence 100 percent, Full
72 Grok 4.20 (Reasoning) xAI 89.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] grok-4.20-beta-0309b-reasoningPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.5 80% confidence 80 percent, High
73 Grok 4.7 xAI 89.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.5 100% confidence 100 percent, Full
74 DeepSeek V4 Flash 0731 DeepSeek 89%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.0 69% confidence 69 percent, Medium
75 Claude Opus 5.5 Anthropic 88.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.5 100% confidence 100 percent, Full
76 DeepSeek V4.1 Flash DeepSeek 88.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.5 69% confidence 69 percent, Medium
77 Claude Opus 4.8 Anthropic 88%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' output effort. [variant] anthropic-opus-4-8-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.0 100% confidence 100 percent, Full
78 GPT-5.6 Luna OpenAI 88%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.0 100% confidence 100 percent, Full
79 GPT-5.6 Luna OpenAI 87.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.7 100% confidence 100 percent, Full
80 Qwen3.8 27B Alibaba / Qwen 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] alibaba-qwen3-8-27b-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.5 69% confidence 69 percent, Medium
81 Grok 4.6 xAI 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] xai-grok-4-6-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.5 100% confidence 100 percent, Full
82 Grok 4.5 xAI 87.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.2 100% confidence 100 percent, Full
83 DeepSeek V4 Pro 0813 DeepSeek 87.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-pro-0813-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.2 69% confidence 69 percent, Medium
84 DeepSeek V4 Flash 0731 DeepSeek 87%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.0 69% confidence 69 percent, Medium
85 Grok 4.6 xAI 87%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] xai-grok-4-6-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.0 100% confidence 100 percent, Full
86 Grok 4.6 xAI 87%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] xai-grok-4-6-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.0 100% confidence 100 percent, Full
87 GPT-6 Luna OpenAI 86.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-luna-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.7 100% confidence 100 percent, Full
88 Kimi K3 Moonshot AI 86.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.7 100% confidence 100 percent, Full
89 Claude Sonnet 4.6 Anthropic 86.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude_sonnet_4_6_highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.5 100% confidence 100 percent, Full
90 GPT-5.2 OpenAI 86.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.2 100% confidence 100 percent, Full
91 GPT-5.4 OpenAI 86.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.2 100% confidence 100 percent, Full
92 Claude Opus 4.6 Anthropic 86%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'low' output effort. [variant] claude-opus-4-6-thinking-120K-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.0 100% confidence 100 percent, Full
93 Claude Sonnet 4.6 Anthropic 86%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude_sonnet_4_6_maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.0 100% confidence 100 percent, Full
94 GPT-5.2 Pro OpenAI 85.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
85.7 48% confidence 48 percent, Low
95 Grok 4.5 xAI 85.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
85.7 100% confidence 100 percent, Full
96 Gemini 3.7 Flash Google 85.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] google-gemini-3-7-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
85.2 100% confidence 100 percent, Full
97 Gemini 3 Flash Preview Google 84.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.7 93% confidence 93 percent, High
98 DeepSeek V4 Flash 0731 DeepSeek 84%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.0 69% confidence 69 percent, Medium
99 Inkling Small Thinking Machines Lab 84%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] thinky-inkling-small-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.0 100% confidence 100 percent, Full
100 GPT-6 Sol OpenAI 83.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
83.7 100% confidence 100 percent, Full
101 Gemini 3.6 Flash Google 83.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-6-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
83.2 100% confidence 100 percent, Full
102 GPT-5.2 Pro OpenAI 81.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
81.2 48% confidence 48 percent, Low
103 Claude Opus 4.5 Anthropic 80%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-64kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
80.0 100% confidence 100 percent, Full
104 Inkling Thinking Machines Lab 79.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] thinky-inklingPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
79.5 100% confidence 100 percent, Full
105 Grok 4.5 xAI 79.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
79.2 100% confidence 100 percent, Full
106 GPT-5.2 OpenAI 78.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
78.7 100% confidence 100 percent, Full
107 Inkling Small Thinking Machines Lab 78%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] thinky-inkling-small-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
78.0 100% confidence 100 percent, Full
108 GPT-5.6 Terra OpenAI 77%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
77.0 100% confidence 100 percent, Full
109 GLM-5.2 Z.ai 77%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5.2Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
77.0 100% confidence 100 percent, Full
110 Gemini 3.6 Flash Google 76.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-6-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
76.5 100% confidence 100 percent, Full
111 GPT-5.6 Luna OpenAI 76.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
76.5 100% confidence 100 percent, Full
112 GPT-5.5 OpenAI 76.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
76.2 100% confidence 100 percent, Full
113 Claude Opus 4.5 Anthropic 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
75.8 100% confidence 100 percent, Full
114 Grok 4.6 xAI 74.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] xai-grok-4-6-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
74.8 100% confidence 100 percent, Full
115 GPT-5.6 Sol OpenAI 74.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
74.5 100% confidence 100 percent, Full
116 GPT-6 Luna OpenAI 73%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
73.0 100% confidence 100 percent, Full
117 GPT-5.1 OpenAI 72.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.8 85% confidence 85 percent, High
118 GPT-5.2 OpenAI 72.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.7 100% confidence 100 percent, Full
119 GPT-6 Sol OpenAI 72.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.2 100% confidence 100 percent, Full
120 Claude Opus 4.5 Anthropic 72%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.0 100% confidence 100 percent, Full
121 GLM-5.3-Flash Z.ai 71.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] zai-glm-5-3-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
71.8 100% confidence 100 percent, Full
122 GPT-6 Luna OpenAI 70.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
70.3 100% confidence 100 percent, Full
123 GPT-5 Pro OpenAI 70.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-pro-2025-10-06Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
70.2 64% confidence 64 percent, Medium
124 Qwen3.8 27B Alibaba / Qwen 69.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] alibaba-qwen3-8-27b-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
69.2 69% confidence 69 percent, Medium
125 Qwen3.8 27B Alibaba / Qwen 68.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] alibaba-qwen3-8-27b-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
68.7 69% confidence 69 percent, Medium
126 GPT-5.4 OpenAI 68.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
68.2 100% confidence 100 percent, Full
127 Inkling Small Thinking Machines Lab 67%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] thinky-inkling-small-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
67.0 100% confidence 100 percent, Full
128 GPT-5 OpenAI 65.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
65.7 100% confidence 100 percent, Full
129 Kimi K3 Moonshot AI 65.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
65.7 100% confidence 100 percent, Full
130 Kimi K2.5 Moonshot AI 65.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] kimi-k2.5Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
65.3 100% confidence 100 percent, Full
131 Claude Sonnet 4.5 (latest) Anthropic 63.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
63.7 43% confidence 43 percent, Low
132 MiniMax-M2.5 MiniMax 63.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] minimax-m2.5Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
63.7 100% confidence 100 percent, Full
133 GPT-5.4 mini OpenAI 63.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
63.7 100% confidence 100 percent, Full
134 Grok 4.7 xAI 63.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] xai-grok-4-7-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
63.3 100% confidence 100 percent, Full
135 GPT-6 Luna OpenAI 61%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
61.0 100% confidence 100 percent, Full
136 o3 OpenAI 60.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
60.8 87% confidence 87 percent, High
137 GPT-5.6 Terra OpenAI 60.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
60.2 100% confidence 100 percent, Full
138 o3-pro OpenAI 59.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
59.3 22% confidence 22 percent, Low
139 Claude Opus 4.5 Anthropic 58.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
58.7 100% confidence 100 percent, Full
140 o4-mini OpenAI 58.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
58.7 87% confidence 87 percent, High
141 GPT-5.4 mini OpenAI 58.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
58.0 100% confidence 100 percent, Full
142 Gemini 3 Flash Preview Google 57.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
57.7 93% confidence 93 percent, High
143 GPT-5.1 OpenAI 57.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
57.7 85% confidence 85 percent, High
144 DeepSeek V3.2 DeepSeek 57.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek-v3.2Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
57.0 56% confidence 56 percent, Medium
145 o3-pro OpenAI 57.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
57.0 22% confidence 22 percent, Low
146 GPT-5.6 Luna OpenAI 56.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
56.5 100% confidence 100 percent, Full
147 GPT-5 OpenAI 56.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
56.2 100% confidence 100 percent, Full
148 GPT-5.2 OpenAI 55.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
55.7 100% confidence 100 percent, Full
149 GPT-5 Mini OpenAI 54.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
54.3 99% confidence 99 percent, High
150 o3 OpenAI 53.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
53.8 87% confidence 87 percent, High
151 Gemini 3.5 Flash Lite Google 53.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-5-flash-lite-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
53.5 100% confidence 100 percent, Full
152 Inkling Small Thinking Machines Lab 52.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] thinky-inkling-small-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
52.5 100% confidence 100 percent, Full
153 GPT-5.4 nano OpenAI 51.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
51.5 100% confidence 100 percent, Full
154 Claude Sonnet 4.5 (latest) Anthropic 48.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
48.3 43% confidence 43 percent, Low
155 Claude Haiku 4.5 (latest) Anthropic 47.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
47.7 33% confidence 33 percent, Low
156 GLM-5.3-Flash Z.ai 47.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] zai-glm-5-3-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
47.0 100% confidence 100 percent, Full
157 Claude Sonnet 4.5 (latest) Anthropic 46.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
46.5 43% confidence 43 percent, Low
158 GLM-5 Z.ai 44.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
44.7 90% confidence 90 percent, High
159 o3-pro OpenAI 44.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
44.3 22% confidence 22 percent, Low
160 GPT-5 OpenAI 44%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
44.0 100% confidence 100 percent, Full
161 o4-mini OpenAI 41.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
41.8 87% confidence 87 percent, High
162 o3 OpenAI 41.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
41.5 87% confidence 87 percent, High
163 Gemini 2.5 Pro Google 41%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
41.0 90% confidence 90 percent, High
164 GPT-5.4 mini OpenAI 40.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
40.8 100% confidence 100 percent, Full
165 Claude Opus 4.5 Anthropic 40%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
40.0 100% confidence 100 percent, Full
166 Claude Sonnet 4 Anthropic 40%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
40.0 100% confidence 100 percent, Full
167 GPT-5.4 nano OpenAI 38.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
38.2 100% confidence 100 percent, Full
168 GPT-6 Luna OpenAI 37.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
37.7 100% confidence 100 percent, Full
169 Claude Haiku 4.5 (latest) Anthropic 37.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
37.3 33% confidence 33 percent, Low
170 GPT-5 Mini OpenAI 37.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
37.3 99% confidence 99 percent, High
171 Gemini 2.5 Pro Google 37%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
37.0 90% confidence 90 percent, High
172 Claude Opus 4 Anthropic 35.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
35.7 100% confidence 100 percent, Full
173 o3-mini OpenAI 34.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
34.5 96% confidence 96 percent, High
174 GPT-5.6 Luna OpenAI 34.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
34.2 100% confidence 100 percent, Full
175 GPT-5.1 OpenAI 33.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
33.2 85% confidence 85 percent, High
176 GPT-5.4 nano OpenAI 33%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
33.0 100% confidence 100 percent, Full
177 Gemini 3.5 Flash Lite Google 32.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-5-flash-lite-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
32.3 100% confidence 100 percent, Full
178 Claude Sonnet 4.5 (latest) Anthropic 31%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
31.0 43% confidence 43 percent, Low
179 Claude Opus 4 Anthropic 30.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
30.7 100% confidence 100 percent, Full
180 Gemini 2.5 Pro Google 29.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
29.5 90% confidence 90 percent, High
181 Claude Sonnet 4 Anthropic 29.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
29.0 100% confidence 100 percent, Full
182 Gemini 3 Flash Preview Google 29.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
29.0 93% confidence 93 percent, High
183 Claude Sonnet 3.7 Anthropic 28.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
28.6 100% confidence 100 percent, Full
184 Claude Sonnet 4 Anthropic 28.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
28.0 100% confidence 100 percent, Full
185 Claude Opus 4 Anthropic 27%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
27.0 100% confidence 100 percent, Full
186 GPT-5 Mini OpenAI 26.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
26.3 99% confidence 99 percent, High
187 Claude Haiku 4.5 (latest) Anthropic 25.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
25.5 33% confidence 33 percent, Low
188 Claude Sonnet 4.5 (latest) Anthropic 25.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
25.5 43% confidence 43 percent, Low
189 Claude Sonnet 4 Anthropic 23.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
23.8 100% confidence 100 percent, Full
190 Claude Opus 4 Anthropic 22.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
22.5 100% confidence 100 percent, Full
191 o3-mini OpenAI 22.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
22.3 96% confidence 96 percent, High
192 o4-mini OpenAI 21.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
21.3 87% confidence 87 percent, High
193 DeepSeek-R1 DeepSeek 21.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
21.2 88% confidence 88 percent, High
194 Claude Sonnet 3.7 Anthropic 21.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 8KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
21.2 100% confidence 100 percent, Full
195 GPT-5 Nano OpenAI 20.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
20.7 85% confidence 85 percent, High
196 GPT-5.4 nano OpenAI 18.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
18.3 100% confidence 100 percent, Full
197 Gemini 3.5 Flash Lite Google 17%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-5-flash-lite-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
17.0 100% confidence 100 percent, Full
198 Claude Haiku 4.5 (latest) Anthropic 16.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
16.8 33% confidence 33 percent, Low
199 GPT-5 Nano OpenAI 16.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
16.7 85% confidence 85 percent, High
200 Gemini 2.5 Pro Google 16%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
16.0 90% confidence 90 percent, High
201 DeepSeek-R1 DeepSeek 15.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] R1Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
15.8 88% confidence 88 percent, High
202 o3-mini OpenAI 14.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
14.5 96% confidence 96 percent, High
203 Claude Haiku 4.5 (latest) Anthropic 14.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
14.3 33% confidence 33 percent, Low
204 Claude Sonnet 3.7 Anthropic 13.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
13.6 100% confidence 100 percent, Full
205 GPT-5.4 mini OpenAI 13%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
13.0 100% confidence 100 percent, Full
206 GPT-5.2 OpenAI 12.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
12.3 100% confidence 100 percent, Full
207 Claude Sonnet 3.7 Anthropic 11.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 1KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
11.6 100% confidence 100 percent, Full
208 Qwen3 235B-A22B Instruct 2507 Alibaba / Qwen 11%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] qwen3-235b-a22b-instruct-2507Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
11.0 48% confidence 48 percent, Low
209 Magistral Medium (latest) Mistral AI 6.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506-thinkingPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
6.1 16% confidence 16 percent, Low
210 Magistral Medium (latest) Mistral AI 5.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.9 16% confidence 16 percent, Low
211 GPT-5.1 OpenAI 5.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.8 85% confidence 85 percent, High
212 GPT-4.1 OpenAI 5.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-2025-04-14Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.5 100% confidence 100 percent, Full
213 Magistral Small Mistral AI 5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.0 48% confidence 48 percent, Low
214 GPT-4o OpenAI 4.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4o-2024-11-20Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.5 59% confidence 59 percent, Medium
215 Llama 4 Maverick 17B Instruct Meta 4.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Llama-4-Maverick-17B-128E-Instruct-FP8-togetherPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.4 78% confidence 78 percent, Medium
216 GPT-5 Nano OpenAI 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.0 85% confidence 85 percent, High
217 GPT-4.1 mini OpenAI 3.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-mini-2025-04-14Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
3.5 87% confidence 87 percent, High
218 GPT-4.1 nano OpenAI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-nano-2025-04-14Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 82% confidence 82 percent, High

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed