ARC-AGI-2Benchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Claude Opus 5 Anthropic | 90.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effortPublished Jul 24, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 90.4 | 100% confidence 100 percent, Full |
| 2 | GPT-5.5 OpenAI | 85%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] VerifiedPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 85.0 | 100% confidence 100 percent, Full |
| 3 | GPT-5.4 Pro OpenAI | 83.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] VerifiedPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 83.3 | 69% confidence 69 percent, Medium |
| 4 | Gemini 3.1 Pro Preview Google | 77.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published May 19, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 77.1 | 100% confidence 100 percent, Full |
| 5 | GPT-5.4 OpenAI | 73.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] VerifiedPublished Apr 23, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 73.3 | 100% confidence 100 percent, Full |
| 6 | Gemini 3.5 Flash Google | 72.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published May 19, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 72.1 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.