BenchmarksAI benchmark definitions, weights and sources

58 benchmarks feed the SI Score, grouped into 4 capability pillars. Each result on the site links back to the board or dataset it came from.

reasoning

Pillar weight 30%
  1. GPQA Diamond 96% GPT-6 Astra 138
  2. ARC-AGI-2 (semi-private) 95% GPT-6 Astra 74
  3. ARC-AGI-1 (semi-private) 98.5% Claude Fable 5 72
  4. ARC-AGI-2 (public eval) 99.2% Claude Fable 5.1 72
  5. ARC-AGI-1 (public eval) 99% Claude Fable 5.1 69
  6. LiveBench Reasoning: connections 100% DeepSeek V4 Pro 0813 61
  7. LiveBench Reasoning: consecutive events 92.7% Claude Sonnet 4.6 61
  8. LiveBench Reasoning: logic with navigation 88% Muse Spark 1.2 61
  9. LiveBench Reasoning: spatial 100% Qwen3.6 27B 61
  10. LiveBench Reasoning: theory of mind 88.5% GPT-5.4 61
  11. LiveBench Reasoning: zebra puzzles 100% Qwen3.8 Max 61
  12. Humanity's Last Exam (Scale AI) 54.8% GPT-6 Astra 28
  13. Humanity's Last Exam (with tools) 67.7% Claude Opus 5.5 26
  14. Humanity's Last Exam 60.9% Claude Fable 5.1 24
  15. ARC-AGI-3 (semi-private) 62.7% GPT-6 Astra 18
  16. ARC-AGI-2 90.4% Claude Opus 5 6
  17. MMLU-Pro 90.3% Sakana Namazu 5
  18. ARC-AGI-1 97.5% Claude Opus 5 2
  19. ARC-AGI-3 99.9% GPT-6 Astra 2
  20. Humanity's Last Exam (full set, text + multimodal) 44.4% Gemini 3.1 Pro Preview 2
  21. Humanity's Last Exam (text only) 37% Hy3 2
  22. Humanity's Last Exam (text-only subset) 40.5% GLM-5.2 2
  23. Humanity's Last Exam (1,811 verified items) 54.9% Gemini 3.8 Flash 1
  24. Humanity's Last Exam (full set) 36.8% DeepSeek V4.1 Flash 1
  25. Humanity's Last Exam (full set, with tools) 55.3% GLM-5.3-Flash 1
  26. Humanity's Last Exam (text-only subset, with tools) 54.7% GLM-5.2 1
  27. Humanity's Last Exam (text only, with tools) 53.2% Hy3 1

math

Pillar weight 15%
  1. OTIS Mock AIME 2024–2025 100% Qwen3.8 Max 0902 124
  2. FrontierMath Tiers 1–3 (v2) 93.7% GPT-6 Astra 77
  3. FrontierMath Tier 4 (v2) 100% GPT-6.1 Sol 61
  4. LiveBench Math: AMPS Hard 99.0% Claude Opus 5 61
  5. LiveBench Math: integrals 100% Claude Sonnet 5.5 61
  6. LiveBench Math: competition math 98.0% Claude Haiku 5.5 61
  7. LiveBench Math: olympiad 93.0% Claude Fable 5.1 61
  8. LiveBench Math: simplify 75.7% Gemini 3.6 Flash 61
  9. MATH Level 5 98.1% GPT-5 35
  10. FrontierMath Tiers 1–3 52.4% GPT-5.5 Pro 4
  11. FrontierMath Tier 4 39.6% GPT-5.5 Pro 4
  12. FrontierMath Tiers 1–3 (v2) 89% GPT-5.6 Sol 3
  13. FrontierMath Tier 4 (v2) 97.6% GPT-6 Astra 1

coding

Pillar weight 40%
  1. SWE-bench Verified 96% Claude Opus 5 91
  2. LiveBench Coding: code completion 87.0% Claude Opus 5.5 61
  3. LiveBench Coding: code generation 93.0% Claude Sonnet 5.5 61
  4. LiveBench Coding: JavaScript 81.8% DeepSeek V4.1 Flash 61
  5. LiveBench Coding: Python 90% DeepSeek V4.1 Flash 61
  6. LiveBench Coding: TypeScript 60% Claude Fable 5.1 61
  7. Terminal-Bench 2.1 90.6% DeepSeek V4.1 Flash 42
  8. SWE-bench Pro 80.3% Claude Fable 5 38
  9. Terminal-Bench 82.1% Fugu Ultra 28
  10. Terminal-Bench 4.0 70.6% Claude Sonnet 5.5 28
  11. Aider Polyglot 84.9% o3-pro 24
  12. SWE-bench Pro (public) 59.1% GPT-5.4 24
  13. Terminal-Bench 2.0 82.7% GPT-5.5 9
  14. Terminal-Bench 0.1 64.6% GPT-6 Astra 3
  15. SWE-bench Pro (Qwen-corrected tasks) 67.7% Qwen3.8 Max 2
  16. Terminal-Bench 3.0 30% DeepSeek V4.1 Flash 2
  17. SWE-bench Pro (system card, Jun 2026) 80.3% Claude Mythos 5 1

preference

Pillar weight 15%
  1. LMArena Text 1514.8 elo Claude Opus 5.5 165

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed