ARC-AGI-3 (semi-private)Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | GPT-6 Astra OpenAI | 62.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] openai-gpt-6-astra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 62.7 | 100% confidence 100 percent, Full |
| 2 | GPT-6 Astra OpenAI | 59.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] openai-gpt-6-astra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 59.3 | 100% confidence 100 percent, Full |
| 3 | GPT-6 Astra OpenAI | 54.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] openai-gpt-6-astra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 54.8 | 100% confidence 100 percent, Full |
| 4 | GPT-6.1 Sol OpenAI | 52.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 52.7 | 100% confidence 100 percent, Full |
| 5 | GPT-6.1 Sol OpenAI | 39.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 39.9 | 100% confidence 100 percent, Full |
| 6 | GPT-6 Astra OpenAI | 38.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort. [variant] openai-gpt-6-astra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 38.6 | 100% confidence 100 percent, Full |
| 7 | Claude Opus 5 Anthropic | 30.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-claude-opus-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 30.2 | 100% confidence 100 percent, Full |
| 8 | GPT-6.1 Sol OpenAI | 26.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 26.7 | 100% confidence 100 percent, Full |
| 9 | GPT-6 Astra OpenAI | 17.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] openai-gpt-6-astra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 17.5 | 100% confidence 100 percent, Full |
| 10 | GPT-6.1 Sol OpenAI | 10.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 10.6 | 100% confidence 100 percent, Full |
| 11 | Gemini 3.8 Flash Google | 10.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] google-gemini-3-8-flash-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 10.4 | 100% confidence 100 percent, Full |
| 12 | GPT-5.6 Sol OpenAI | 7.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.8 | 100% confidence 100 percent, Full |
| 13 | GPT-5.6 Sol OpenAI | 7.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 7.0 | 100% confidence 100 percent, Full |
| 14 | Gemini 3.8 Flash Google | 6.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 6.0 | 100% confidence 100 percent, Full |
| 15 | GPT-6 Sol OpenAI | 4.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. [variant] openai-gpt-6-sol-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.6 | 100% confidence 100 percent, Full |
| 16 | Gemini 3.8 Flash Google | 4.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] google-gemini-3-8-flash-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 4.0 | 100% confidence 100 percent, Full |
| 17 | GPT-6.1 Sol OpenAI | 3.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-1-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 3.9 | 100% confidence 100 percent, Full |
| 18 | GPT-5.6 Sol OpenAI | 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.1 | 100% confidence 100 percent, Full |
| 19 | Grok 4.6 xAI | 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort. [variant] xai-grok-4-6-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 2.1 | 100% confidence 100 percent, Full |
| 20 | Grok 4.7 xAI | 1.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.8 | 100% confidence 100 percent, Full |
| 21 | GPT-6 Sol OpenAI | 1.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.8 | 100% confidence 100 percent, Full |
| 22 | Grok 4.7 xAI | 1.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.7 | 100% confidence 100 percent, Full |
| 23 | Grok 4.7 xAI | 1.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] xai-grok-4-7-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.7 | 100% confidence 100 percent, Full |
| 24 | Claude Opus 4.8 Anthropic | 1.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' output effort. [variant] anthropic-opus-4-8-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.5 | 100% confidence 100 percent, Full |
| 25 | GPT-5.6 Sol OpenAI | 1.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 1.1 | 100% confidence 100 percent, Full |
| 26 | GPT-6 Sol OpenAI | 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.9 | 100% confidence 100 percent, Full |
| 27 | GPT-5.6 Terra OpenAI | 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.8 | 100% confidence 100 percent, Full |
| 28 | GPT-5.6 Terra OpenAI | 0.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.7 | 100% confidence 100 percent, Full |
| 29 | GPT-5.6 Terra OpenAI | 0.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.5 | 100% confidence 100 percent, Full |
| 30 | GPT-5.5 OpenAI | 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-5-2026-04-23-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.4 | 100% confidence 100 percent, Full |
| 31 | Gemini 3.1 Pro Preview Google | 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] google-gemini-3-1-pro-previewPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.4 | 100% confidence 100 percent, Full |
| 32 | GPT-6 Sol OpenAI | 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.4 | 100% confidence 100 percent, Full |
| 33 | GPT-5.6 Sol OpenAI | 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.3 | 100% confidence 100 percent, Full |
| 34 | Grok 4.5 xAI | 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.3 | 100% confidence 100 percent, Full |
| 35 | Grok 4.5 xAI | 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.3 | 100% confidence 100 percent, Full |
| 36 | Grok 4.5 xAI | 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-5-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.3 | 100% confidence 100 percent, Full |
| 37 | Grok 4.7 xAI | 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] xai-grok-4-7-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.2 | 100% confidence 100 percent, Full |
| 38 | GPT-5.4 OpenAI | 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-4-2026-03-05-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.2 | 100% confidence 100 percent, Full |
| 39 | GPT-6 Luna OpenAI | 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'medium' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.2 | 100% confidence 100 percent, Full |
| 40 | GPT-6 Luna OpenAI | 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.2 | 100% confidence 100 percent, Full |
| 41 | Claude Opus 4.7 Anthropic | 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] anthropic-opus-4-7-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.2 | 100% confidence 100 percent, Full |
| 42 | GPT-5.6 Luna OpenAI | 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.2 | 100% confidence 100 percent, Full |
| 43 | GPT-5.6 Luna OpenAI | 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.2 | 100% confidence 100 percent, Full |
| 44 | GPT-5.6 Luna OpenAI | 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.2 | 100% confidence 100 percent, Full |
| 45 | GPT-6 Luna OpenAI | 0.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'xhigh' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.2 | 100% confidence 100 percent, Full |
| 46 | GPT-6 Sol OpenAI | 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-sol-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.1 | 100% confidence 100 percent, Full |
| 47 | GPT-6 Luna OpenAI | 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort using the standard harness. Partial dataset coverage: scores use full catalog denominators; costs include recorded usage only, excluding unrecorded attempts. [variant] openai-gpt-6-luna-maxPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.1 | 100% confidence 100 percent, Full |
| 48 | GPT-5.6 Luna OpenAI | 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-highPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.1 | 100% confidence 100 percent, Full |
| 49 | Grok 4.20 (Reasoning) xAI | 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] xai-grok-4-20-beta-0309-reasoningPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.1 | 80% confidence 80 percent, High |
| 50 | GPT-5.6 Terra OpenAI | 0.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-mediumPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.1 | 100% confidence 100 percent, Full |
| 51 | GPT-6 Luna OpenAI | 0.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort using the standard harness. [variant] openai-gpt-6-luna-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 52 | GPT-5.6 Luna OpenAI | 0.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-luna-xhighPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
| 53 | GPT-5.6 Terra OpenAI | 0.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] openai-gpt-5-6-terra-lowPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation Open source ↗ | 0.0 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.