BenchmarksAI benchmark definitions, weights and sources
58 benchmarks feed the SI Score, grouped into 4 capability pillars. Each result on the site links back to the board or dataset it came from.
reasoning
Pillar weight 30%- GPQA Diamond 96% GPT-6 Astra 138
- ARC-AGI-2 (semi-private) 95% GPT-6 Astra 74
- ARC-AGI-1 (semi-private) 98.5% Claude Fable 5 72
- ARC-AGI-2 (public eval) 99.2% Claude Fable 5.1 72
- ARC-AGI-1 (public eval) 99% Claude Fable 5.1 69
- LiveBench Reasoning: connections 100% DeepSeek V4 Pro 0813 61
- LiveBench Reasoning: consecutive events 92.7% Claude Sonnet 4.6 61
- LiveBench Reasoning: logic with navigation 88% Muse Spark 1.2 61
- LiveBench Reasoning: spatial 100% Qwen3.6 27B 61
- LiveBench Reasoning: theory of mind 88.5% GPT-5.4 61
- LiveBench Reasoning: zebra puzzles 100% Qwen3.8 Max 61
- Humanity's Last Exam (Scale AI) 54.8% GPT-6 Astra 28
- Humanity's Last Exam (with tools) 67.7% Claude Opus 5.5 26
- Humanity's Last Exam 60.9% Claude Fable 5.1 24
- ARC-AGI-3 (semi-private) 62.7% GPT-6 Astra 18
- ARC-AGI-2 90.4% Claude Opus 5 6
- MMLU-Pro 90.3% Sakana Namazu 5
- ARC-AGI-1 97.5% Claude Opus 5 2
- ARC-AGI-3 99.9% GPT-6 Astra 2
- Humanity's Last Exam (full set, text + multimodal) 44.4% Gemini 3.1 Pro Preview 2
- Humanity's Last Exam (text only) 37% Hy3 2
- Humanity's Last Exam (text-only subset) 40.5% GLM-5.2 2
- Humanity's Last Exam (1,811 verified items) 54.9% Gemini 3.8 Flash 1
- Humanity's Last Exam (full set) 36.8% DeepSeek V4.1 Flash 1
- Humanity's Last Exam (full set, with tools) 55.3% GLM-5.3-Flash 1
- Humanity's Last Exam (text-only subset, with tools) 54.7% GLM-5.2 1
- Humanity's Last Exam (text only, with tools) 53.2% Hy3 1
math
Pillar weight 15%- OTIS Mock AIME 2024–2025 100% Qwen3.8 Max 0902 124
- FrontierMath Tiers 1–3 (v2) 93.7% GPT-6 Astra 77
- FrontierMath Tier 4 (v2) 100% GPT-6.1 Sol 61
- LiveBench Math: AMPS Hard 99.0% Claude Opus 5 61
- LiveBench Math: integrals 100% Claude Sonnet 5.5 61
- LiveBench Math: competition math 98.0% Claude Haiku 5.5 61
- LiveBench Math: olympiad 93.0% Claude Fable 5.1 61
- LiveBench Math: simplify 75.7% Gemini 3.6 Flash 61
- MATH Level 5 98.1% GPT-5 35
- FrontierMath Tiers 1–3 52.4% GPT-5.5 Pro 4
- FrontierMath Tier 4 39.6% GPT-5.5 Pro 4
- FrontierMath Tiers 1–3 (v2) 89% GPT-5.6 Sol 3
- FrontierMath Tier 4 (v2) 97.6% GPT-6 Astra 1
coding
Pillar weight 40%- SWE-bench Verified 96% Claude Opus 5 91
- LiveBench Coding: code completion 87.0% Claude Opus 5.5 61
- LiveBench Coding: code generation 93.0% Claude Sonnet 5.5 61
- LiveBench Coding: JavaScript 81.8% DeepSeek V4.1 Flash 61
- LiveBench Coding: Python 90% DeepSeek V4.1 Flash 61
- LiveBench Coding: TypeScript 60% Claude Fable 5.1 61
- Terminal-Bench 2.1 90.6% DeepSeek V4.1 Flash 42
- SWE-bench Pro 80.3% Claude Fable 5 38
- Terminal-Bench 82.1% Fugu Ultra 28
- Terminal-Bench 4.0 70.6% Claude Sonnet 5.5 28
- Aider Polyglot 84.9% o3-pro 24
- SWE-bench Pro (public) 59.1% GPT-5.4 24
- Terminal-Bench 2.0 82.7% GPT-5.5 9
- Terminal-Bench 0.1 64.6% GPT-6 Astra 3
- SWE-bench Pro (Qwen-corrected tasks) 67.7% Qwen3.8 Max 2
- Terminal-Bench 3.0 30% DeepSeek V4.1 Flash 2
- SWE-bench Pro (system card, Jun 2026) 80.3% Claude Mythos 5 1
preference
Pillar weight 15%- LMArena Text 1514.8 elo Claude Opus 5.5 165