ARC-AGI-2 (semi-private)Benchmark scores and sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 1 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 GPT-6 Astra OpenAI 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
95.0 100% confidence 100 percent, Full
2 GPT-6.1 Sol OpenAI 94.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.2 100% confidence 100 percent, Full
3 Claude Opus 5.5 Anthropic 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
93.3 100% confidence 100 percent, Full
4 GPT-6 Astra OpenAI 93.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
93.3 100% confidence 100 percent, Full
5 Claude Opus 5.5 Anthropic 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 100% confidence 100 percent, Full
6 GPT-5.6 Sol OpenAI 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.5 100% confidence 100 percent, Full
7 GPT-6 Astra OpenAI 92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.1 100% confidence 100 percent, Full
8 GPT-6 Astra OpenAI 92.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
92.1 100% confidence 100 percent, Full
9 Claude Opus 5.5 Anthropic 91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] anthropic-claude-opus-5-5-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.7 100% confidence 100 percent, Full
10 GPT-6.1 Sol OpenAI 91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.7 100% confidence 100 percent, Full
11 GPT-6.1 Sol OpenAI 91.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
91.7 100% confidence 100 percent, Full
12 Claude Opus 5 Anthropic 90.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' output effort. [variant] anthropic-claude-opus-5-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.4 100% confidence 100 percent, Full
13 Claude Fable 5.1 Anthropic 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-1-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.0 100% confidence 100 percent, Full
14 Claude Fable 5.1 Anthropic 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-1-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.0 100% confidence 100 percent, Full
15 GPT-5.6 Sol OpenAI 90%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.0 100% confidence 100 percent, Full
16 GPT-6 Sol OpenAI 89.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-sol-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.6 100% confidence 100 percent, Full
17 Claude Fable 5 Anthropic 89.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] anthropic-claude-fable-5-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.2 100% confidence 100 percent, Full
18 Gemini 3.8 Flash Google 89.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.2 100% confidence 100 percent, Full
19 Claude Fable 5.1 Anthropic 88.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-1-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.8 100% confidence 100 percent, Full
20 Claude Fable 5 Anthropic 88.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] anthropic-claude-fable-5-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.3 100% confidence 100 percent, Full
21 Claude Opus 5 Anthropic 88.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
88.3 100% confidence 100 percent, Full
22 Claude Fable 5 Anthropic 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] anthropic-claude-fable-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.5 100% confidence 100 percent, Full
23 Claude Opus 5.5 Anthropic 87.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
87.5 100% confidence 100 percent, Full
24 GPT-6.1 Sol OpenAI 86.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.7 100% confidence 100 percent, Full
25 Claude Fable 5.1 Anthropic 86.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-1-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
86.3 100% confidence 100 percent, Full
26 GPT-5.6 Sol OpenAI 85.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
85.4 100% confidence 100 percent, Full
27 GPT-6 Astra OpenAI 85.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
85.4 100% confidence 100 percent, Full
28 GPT-5.5 OpenAI 85%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
85.0 100% confidence 100 percent, Full
29 Gemini 3.7 Flash Google 84.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] google-gemini-3-7-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.6 100% confidence 100 percent, Full
30 GPT-5.5 Pro OpenAI 84.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.6 53% confidence 53 percent, Medium
31 GPT-5.5 Pro OpenAI 84.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.2 53% confidence 53 percent, Medium
32 GPT-5.6 Terra OpenAI 83.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
83.9 100% confidence 100 percent, Full
33 GPT-5.4 Pro OpenAI 83.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-pro-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
83.3 69% confidence 69 percent, Medium
34 GPT-5.5 OpenAI 83.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
83.3 100% confidence 100 percent, Full
35 Gemini 3.8 Flash Google 82.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
82.9 100% confidence 100 percent, Full
36 Claude Fable 5 Anthropic 82.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] anthropic-claude-fable-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
82.5 100% confidence 100 percent, Full
37 Claude Fable 5.1 Anthropic 78.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-1-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
78.3 100% confidence 100 percent, Full
38 GPT-6 Sol OpenAI 78.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
78.1 100% confidence 100 percent, Full
39 Gemini 3.8 Flash Google 77.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
77.5 100% confidence 100 percent, Full
40 Gemini 3.1 Pro Preview Google 77.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-1-pro-previewPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
77.1 100% confidence 100 percent, Full
41 Claude Fable 5 Anthropic 76.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] anthropic-claude-fable-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
76.8 100% confidence 100 percent, Full
42 GPT-6.1 Sol OpenAI 76.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
76.7 100% confidence 100 percent, Full
43 Claude Opus 4.7 Anthropic 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
75.8 100% confidence 100 percent, Full
44 GPT-5.6 Terra OpenAI 74.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
74.2 100% confidence 100 percent, Full
45 GPT-5.4 OpenAI 74.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
74.0 100% confidence 100 percent, Full
46 DeepSeek V4.1 Flash DeepSeek 72.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.9 69% confidence 69 percent, Medium
47 Claude Opus 4.8 Anthropic 72.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-opus-4-8-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.1 100% confidence 100 percent, Full
48 Gemini 3.5 Flash Google 72.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-5-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.1 100% confidence 100 percent, Full
49 Claude Opus 4.8 Anthropic 71.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' output effort. [variant] anthropic-opus-4-8-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
71.7 100% confidence 100 percent, Full
50 GPT-5.5 OpenAI 70.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
70.4 100% confidence 100 percent, Full
51 Claude Opus 5.5 Anthropic 70.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] anthropic-claude-opus-5-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
70.1 100% confidence 100 percent, Full
52 Claude Opus 4.6 Anthropic 69.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude-opus-4-6-thinking-120K-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
69.2 100% confidence 100 percent, Full
53 GPT-6 Sol OpenAI 68.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
68.9 100% confidence 100 percent, Full
54 Claude Opus 4.6 Anthropic 68.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude-opus-4-6-thinking-120K-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
68.8 100% confidence 100 percent, Full
55 Claude Opus 4.7 Anthropic 68.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
68.3 100% confidence 100 percent, Full
56 Claude Opus 4.7 Anthropic 67.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
67.5 100% confidence 100 percent, Full
57 DeepSeek V4.1 Flash DeepSeek 67.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] deepseek-v4-1-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
67.5 69% confidence 69 percent, Medium
58 GPT-5.4 OpenAI 67.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
67.5 100% confidence 100 percent, Full
59 Grok 4.6 xAI 67.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] xai-grok-4-6-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
67.1 100% confidence 100 percent, Full
60 GPT-5.6 Sol OpenAI 67.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
67.1 100% confidence 100 percent, Full
61 GPT-5.6 Terra OpenAI 67.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
67.1 100% confidence 100 percent, Full
62 Claude Opus 4.6 Anthropic 66.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'medium' output effort. [variant] claude-opus-4-6-thinking-120K-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
66.3 100% confidence 100 percent, Full
63 GLM-5.3-Flash Z.ai 65.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] zai-glm-5-3-flash-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
65.8 100% confidence 100 percent, Full
64 Grok 4.20 (Reasoning) xAI 65.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] grok-4.20-beta-0309b-reasoningPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
65.1 80% confidence 80 percent, High
65 Grok 4.6 xAI 65.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] xai-grok-4-6-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
65.1 100% confidence 100 percent, Full
66 Claude Opus 4.6 Anthropic 64.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'low' output effort. [variant] claude-opus-4-6-thinking-120K-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
64.6 100% confidence 100 percent, Full
67 Gemini 3.7 Flash Google 63.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] google-gemini-3-7-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
63.7 100% confidence 100 percent, Full
68 Claude Opus 4.8 Anthropic 62.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' output effort. [variant] anthropic-opus-4-8-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
62.2 100% confidence 100 percent, Full
69 Claude Opus 4.7 Anthropic 62.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-4-7-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
62.1 100% confidence 100 percent, Full
70 DeepSeek V4 Flash 0731 DeepSeek 61.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
61.4 69% confidence 69 percent, Medium
71 Grok 4.7 xAI 61.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
61.4 100% confidence 100 percent, Full
72 DeepSeek V4 Pro 0813 DeepSeek 61.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-pro-0813-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
61.3 69% confidence 69 percent, Medium
73 Grok 4.6 xAI 61.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] xai-grok-4-6-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
61.3 100% confidence 100 percent, Full
74 DeepSeek V4.1 Flash DeepSeek 60.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] deepseek-v4-1-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
60.6 69% confidence 69 percent, Medium
75 Claude Sonnet 4.6 Anthropic 60.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'high' output effort. [variant] claude_sonnet_4_6_highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
60.4 100% confidence 100 percent, Full
76 Gemini 3.6 Flash Google 60.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-6-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
60.4 100% confidence 100 percent, Full
77 Kimi K3 Moonshot AI 60.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
60.4 100% confidence 100 percent, Full
78 DeepSeek V4 Pro 0813 DeepSeek 59.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-pro-0813-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
59.7 69% confidence 69 percent, Medium
79 GPT-5.6 Luna OpenAI 59.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
59.5 100% confidence 100 percent, Full
80 GPT-6 Luna OpenAI 59.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-luna-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
59.3 100% confidence 100 percent, Full
81 Grok 4.7 xAI 58.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
58.8 100% confidence 100 percent, Full
82 Grok 4.7 xAI 58.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
58.3 100% confidence 100 percent, Full
83 Claude Sonnet 4.6 Anthropic 58.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with a 120K thinking budget and a maximum context window of 128K tokens and 'max' output effort. [variant] claude_sonnet_4_6_maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
58.3 100% confidence 100 percent, Full
84 GPT-6 Sol OpenAI 57.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
57.8 100% confidence 100 percent, Full
85 DeepSeek V4 Pro 0813 DeepSeek 56.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-pro-0813-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
56.3 69% confidence 69 percent, Medium
86 DeepSeek V4 Flash 0731 DeepSeek 56.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
56.0 69% confidence 69 percent, Medium
87 GPT-5.4 OpenAI 55.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
55.4 100% confidence 100 percent, Full
88 Kimi K3 Moonshot AI 55.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
55.0 100% confidence 100 percent, Full
89 GPT-5.2 Pro OpenAI 54.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
54.2 48% confidence 48 percent, Low
90 Gemini 3.7 Flash Google 52.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] google-gemini-3-7-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
52.9 100% confidence 100 percent, Full
91 GPT-5.2 OpenAI 52.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
52.9 100% confidence 100 percent, Full
92 Grok 4.5 xAI 52.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
52.6 100% confidence 100 percent, Full
93 Grok 4.5 xAI 52.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
52.6 100% confidence 100 percent, Full
94 Gemini 3.6 Flash Google 50.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-6-flash-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
50.4 100% confidence 100 percent, Full
95 GLM-5.3-Flash Z.ai 50.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] zai-glm-5-3-flash-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
50.1 100% confidence 100 percent, Full
96 GPT-5.6 Luna OpenAI 47.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
47.6 100% confidence 100 percent, Full
97 DeepSeek V4 Flash 0731 DeepSeek 46.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
46.0 69% confidence 69 percent, Medium
98 GPT-5.2 OpenAI 43.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
43.3 100% confidence 100 percent, Full
99 GPT-5.6 Sol OpenAI 42.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
42.5 100% confidence 100 percent, Full
100 Qwen3.8 27B Alibaba / Qwen 42.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] alibaba-qwen3-8-27b-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
42.4 69% confidence 69 percent, Medium
101 GPT-6 Luna OpenAI 41.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
41.9 100% confidence 100 percent, Full
102 Inkling Small Thinking Machines Lab 40.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] thinky-inkling-small-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
40.1 100% confidence 100 percent, Full
103 GPT-5.2 Pro OpenAI 38.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-pro-2025-12-11-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
38.5 48% confidence 48 percent, Low
104 Claude Opus 4.5 Anthropic 37.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-64kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
37.6 100% confidence 100 percent, Full
105 GPT-5.6 Terra OpenAI 37.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
37.5 100% confidence 100 percent, Full
106 Inkling Thinking Machines Lab 36.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] thinky-inklingPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
36.5 100% confidence 100 percent, Full
107 Gemini 3 Flash Preview Google 33.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
33.6 93% confidence 93 percent, High
108 GPT-5.5 OpenAI 33.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-2026-04-22-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
33.3 100% confidence 100 percent, Full
109 Grok 4.5 xAI 33.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
33.1 100% confidence 100 percent, Full
110 Inkling Small Thinking Machines Lab 33.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] thinky-inkling-small-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
33.1 100% confidence 100 percent, Full
111 GPT-6 Sol OpenAI 31.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
31.5 100% confidence 100 percent, Full
112 GPT-6 Luna OpenAI 31.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
31.4 100% confidence 100 percent, Full
113 Gemini 3.6 Flash Google 30.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-6-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
30.4 100% confidence 100 percent, Full
114 GPT-5.6 Luna OpenAI 29.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
29.3 100% confidence 100 percent, Full
115 GPT-5.4 OpenAI 29.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
29.2 100% confidence 100 percent, Full
116 GLM-5.3-Flash Z.ai 27.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] zai-glm-5-3-flash-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
27.9 100% confidence 100 percent, Full
117 Grok 4.6 xAI 27.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] xai-grok-4-6-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
27.6 100% confidence 100 percent, Full
118 GPT-5.2 OpenAI 26.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
26.7 100% confidence 100 percent, Full
119 Claude Opus 4.5 Anthropic 22.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
22.8 100% confidence 100 percent, Full
120 Qwen3.8 27B Alibaba / Qwen 22.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] alibaba-qwen3-8-27b-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
22.8 69% confidence 69 percent, Medium
121 GLM-5.2 Z.ai 22.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5.2Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
22.8 100% confidence 100 percent, Full
122 GPT-5.4 mini OpenAI 18.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
18.9 100% confidence 100 percent, Full
123 GPT-5.6 Terra OpenAI 18.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
18.8 100% confidence 100 percent, Full
124 GPT-5 Pro OpenAI 18.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-pro-2025-10-06Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
18.3 64% confidence 64 percent, Medium
125 GPT-6 Luna OpenAI 18.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
18.1 100% confidence 100 percent, Full
126 GPT-5.1 OpenAI 17.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
17.6 85% confidence 85 percent, High
127 Grok 4.7 xAI 16.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] xai-grok-4-7-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
16.7 100% confidence 100 percent, Full
128 Claude Opus 4.5 Anthropic 13.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
13.9 100% confidence 100 percent, Full
129 Inkling Small Thinking Machines Lab 13.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] thinky-inkling-small-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
13.6 100% confidence 100 percent, Full
130 Claude Sonnet 4.5 (latest) Anthropic 13.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
13.6 43% confidence 43 percent, Low
131 Qwen3.8 27B Alibaba / Qwen 13.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] alibaba-qwen3-8-27b-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
13.2 69% confidence 69 percent, Medium
132 GPT-5.4 mini OpenAI 13.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
13.2 100% confidence 100 percent, Full
133 Gemini 3 Flash Preview Google 12.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
12.8 93% confidence 93 percent, High
134 Kimi K3 Moonshot AI 12.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
12.4 100% confidence 100 percent, Full
135 Kimi K2.5 Moonshot AI 11.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] kimi-k2.5Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
11.8 100% confidence 100 percent, Full
136 Gemini 3.5 Flash Lite Google 10.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] gemini-3-5-flash-lite-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
10.3 100% confidence 100 percent, Full
137 GPT-5 OpenAI 9.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
9.9 100% confidence 100 percent, Full
138 GPT-5.2 OpenAI 9.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
9.7 100% confidence 100 percent, Full
139 Claude Opus 4 Anthropic 8.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
8.6 100% confidence 100 percent, Full
140 Claude Opus 4.5 Anthropic 7.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-5-20251101-thinking-nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
7.8 100% confidence 100 percent, Full
141 GPT-5 OpenAI 7.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
7.5 100% confidence 100 percent, Full
142 GPT-5.6 Luna OpenAI 7.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
7.4 100% confidence 100 percent, Full
143 Claude Sonnet 4.5 (latest) Anthropic 6.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
6.9 43% confidence 43 percent, Low
144 Claude Sonnet 4.5 (latest) Anthropic 6.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
6.9 43% confidence 43 percent, Low
145 GPT-5.1 OpenAI 6.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
6.5 85% confidence 85 percent, High
146 o3 OpenAI 6.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
6.5 87% confidence 87 percent, High
147 o4-mini OpenAI 6.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
6.1 87% confidence 87 percent, High
148 Claude Sonnet 4 Anthropic 5.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.9 100% confidence 100 percent, Full
149 Claude Sonnet 4.5 (latest) Anthropic 5.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.8 43% confidence 43 percent, Low
150 GPT-5.4 nano OpenAI 5.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.7 100% confidence 100 percent, Full
151 Gemini 3.5 Flash Lite Google 5.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] gemini-3-5-flash-lite-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.3 100% confidence 100 percent, Full
152 GPT-5.6 Luna OpenAI 5.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.1 100% confidence 100 percent, Full
153 Inkling Small Thinking Machines Lab 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] thinky-inkling-small-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.9 100% confidence 100 percent, Full
154 Gemini 2.5 Pro Google 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.9 90% confidence 90 percent, High
155 MiniMax-M2.5 MiniMax 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] minimax-m2.5Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.9 100% confidence 100 percent, Full
156 o3-pro OpenAI 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.9 22% confidence 22 percent, Low
157 GLM-5 Z.ai 4.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] glm-5Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.9 90% confidence 90 percent, High
158 GPT-6 Luna OpenAI 4.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.6 100% confidence 100 percent, Full
159 Claude Opus 4 Anthropic 4.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.5 100% confidence 100 percent, Full
160 GPT-5 Mini OpenAI 4.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.4 99% confidence 99 percent, High
161 GPT-5.4 mini OpenAI 4.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.4 100% confidence 100 percent, Full
162 Claude Haiku 4.5 (latest) Anthropic 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-32kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.0 33% confidence 33 percent, Low
163 DeepSeek V3.2 DeepSeek 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek-v3.2Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.0 56% confidence 56 percent, Medium
164 Gemini 2.5 Pro Google 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.0 90% confidence 90 percent, High
165 GPT-5 Mini OpenAI 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.0 99% confidence 99 percent, High
166 Claude Sonnet 4.5 (latest) Anthropic 3.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-5-20250929Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
3.8 43% confidence 43 percent, Low
167 GPT-5.4 nano OpenAI 3.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
3.6 100% confidence 100 percent, Full
168 o3-mini OpenAI 3.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
3.0 96% confidence 96 percent, High
169 o3 OpenAI 3.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
3.0 87% confidence 87 percent, High
170 Gemini 2.5 Pro Google 2.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.9 90% confidence 90 percent, High
171 Claude Haiku 4.5 (latest) Anthropic 2.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-16kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.8 33% confidence 33 percent, Low
172 GPT-5 Nano OpenAI 2.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.6 85% confidence 85 percent, High
173 o4-mini OpenAI 2.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.4 87% confidence 87 percent, High
174 Claude Sonnet 4 Anthropic 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.1 100% confidence 100 percent, Full
175 o3-mini OpenAI 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.1 96% confidence 96 percent, High
176 o3-pro OpenAI 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.1 22% confidence 22 percent, Low
177 o3 OpenAI 2.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-2025-04-16-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
2.0 87% confidence 87 percent, High
178 GPT-5 OpenAI 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2025-08-07-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.9 100% confidence 100 percent, Full
179 GPT-5.1 OpenAI 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.9 85% confidence 85 percent, High
180 GPT-5.4 nano OpenAI 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.9 100% confidence 100 percent, Full
181 o3-pro OpenAI 1.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-pro-2025-06-10-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.9 22% confidence 22 percent, Low
182 Claude Haiku 4.5 (latest) Anthropic 1.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-8kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.7 33% confidence 33 percent, Low
183 o4-mini OpenAI 1.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o4-mini-2025-04-16-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.7 87% confidence 87 percent, High
184 GPT-5.4 nano OpenAI 1.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-nano-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.5 100% confidence 100 percent, Full
185 Gemini 3.5 Flash Lite Google 1.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] gemini-3-5-flash-lite-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.5 100% confidence 100 percent, Full
186 DeepSeek-R1 DeepSeek 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] R1Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.3 88% confidence 88 percent, High
187 Gemini 2.0 Flash Google 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Gemini 2.0 FlashPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.3 28% confidence 28 percent, Low
188 Claude Opus 4 Anthropic 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.3 100% confidence 100 percent, Full
189 Claude Sonnet 4 Anthropic 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.3 100% confidence 100 percent, Full
190 Qwen3 235B-A22B Instruct 2507 Alibaba / Qwen 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] qwen3-235b-a22b-instruct-2507Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.3 48% confidence 48 percent, Low
191 Claude Haiku 4.5 (latest) Anthropic 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.3 33% confidence 33 percent, Low
192 Claude Haiku 4.5 (latest) Anthropic 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-haiku-4-5-20251001-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.3 33% confidence 33 percent, Low
193 Gemini 3 Flash Preview Google 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-3-flash-preview-thinking-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.3 93% confidence 93 percent, High
194 DeepSeek-R1 DeepSeek 1.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.1 88% confidence 88 percent, High
195 GPT-5.4 mini OpenAI 1.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-4-mini-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.1 100% confidence 100 percent, Full
196 Claude Sonnet 3.7 Anthropic 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 8KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.9 100% confidence 100 percent, Full
197 GPT-5 Nano OpenAI 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-mediumPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.9 85% confidence 85 percent, High
198 Claude Sonnet 4 Anthropic 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.9 100% confidence 100 percent, Full
199 GPT-5 Mini OpenAI 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-mini-2025-08-07-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.8 99% confidence 99 percent, High
200 GPT-5.2 OpenAI 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-2-2025-12-11-thinking-nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.8 100% confidence 100 percent, Full
201 Claude Sonnet 3.7 Anthropic 0.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.7 100% confidence 100 percent, Full
202 GPT-4.1 OpenAI 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-2025-04-14Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.4 100% confidence 100 percent, Full
203 GPT-5.1 OpenAI 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-1-2025-11-13-thinking-nonePublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.4 85% confidence 85 percent, High
204 Claude Sonnet 3.7 Anthropic 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 1KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.4 100% confidence 100 percent, Full
205 Claude Sonnet 3.7 Anthropic 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 100% confidence 100 percent, Full
206 Claude Opus 4 Anthropic 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-opus-4-20250514-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 100% confidence 100 percent, Full
207 Gemini 2.5 Pro Google 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gemini-2-5-pro-2025-06-17-thinking-1kPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 90% confidence 90 percent, High
208 Llama 4 Maverick 17B Instruct Meta 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Llama-4-Maverick-17B-128E-Instruct-FP8-togetherPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 78% confidence 78 percent, Medium
209 Magistral Medium (latest) Mistral AI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 16% confidence 16 percent, Low
210 Magistral Medium (latest) Mistral AI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-medium-2506-thinkingPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 16% confidence 16 percent, Low
211 Magistral Small Mistral AI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 48% confidence 48 percent, Low
212 GPT-4.1 mini OpenAI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-mini-2025-04-14Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 87% confidence 87 percent, High
213 GPT-4.1 nano OpenAI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4-1-nano-2025-04-14Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 82% confidence 82 percent, High
214 GPT-4o OpenAI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4o-2024-11-20Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 59% confidence 59 percent, Medium
215 GPT-4o mini OpenAI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-4o-mini-2024-07-18Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 100% confidence 100 percent, Full
216 GPT-5 Nano OpenAI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-nano-2025-08-07-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 85% confidence 85 percent, High
217 o3-mini OpenAI 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] o3-mini-2025-01-31-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0 96% confidence 96 percent, High

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed