The SI Score, in plain words
The SI Score is our composite ranking of Super Intelligence models, on a 0–100 scale. We do not run our own evaluations. Instead, we collect 58 published benchmarks from open sources, group them into 4 capability pillars, and combine them with published weights:
| Pillar | Benchmarks | Weight in SI Score |
|---|---|---|
| reasoning | ARC-AGI-1ARC-AGI-2ARC-AGI-3ARC-AGI-1 (public eval)ARC-AGI-1 (semi-private)ARC-AGI-2 (public eval)ARC-AGI-2 (semi-private)ARC-AGI-3 (semi-private)GPQA Diamond ×0.5Humanity's Last ExamHumanity's Last Exam (1,811 verified items)Humanity's Last Exam (full set)Humanity's Last Exam (full set, text + multimodal)Humanity's Last Exam (full set, with tools)Humanity's Last Exam (Scale AI)Humanity's Last Exam (text only)Humanity's Last Exam (text-only subset)Humanity's Last Exam (text-only subset, with tools)Humanity's Last Exam (text only, with tools)Humanity's Last Exam (with tools)LiveBench Reasoning: connectionsLiveBench Reasoning: consecutive eventsLiveBench Reasoning: logic with navigationLiveBench Reasoning: spatialLiveBench Reasoning: theory of mindLiveBench Reasoning: zebra puzzlesMMLU-Pro ×0.5 | 30% |
| math | FrontierMath Tiers 1–3FrontierMath Tier 4FrontierMath Tier 4 (v2)FrontierMath Tiers 1–3 (v2)FrontierMath Tiers 1–3 (v2)FrontierMath Tier 4 (v2)LiveBench Math: AMPS HardLiveBench Math: integralsLiveBench Math: competition mathLiveBench Math: olympiadLiveBench Math: simplifyMATH Level 5OTIS Mock AIME 2024–2025 ×0.5 | 15% |
| coding | Aider PolyglotLiveBench Coding: code completionLiveBench Coding: code generationLiveBench Coding: JavaScriptLiveBench Coding: PythonLiveBench Coding: TypeScriptSWE-bench ProSWE-bench Pro (public)SWE-bench Pro (Qwen-corrected tasks)SWE-bench Pro (system card, Jun 2026)SWE-bench VerifiedTerminal-BenchTerminal-Bench 0.1Terminal-Bench 2.0Terminal-Bench 2.1Terminal-Bench 3.0Terminal-Bench 4.0 | 40% |
| preference | LMArena Text | 15% |
The production method si-v3-retained-evidence-2 uses fixed absolute scales, not percentiles.
Percentage results retain their 0–100 value (inverted for lower-is-better metrics). Elo is mapped
through a logistic curve centered at 1,200 with scale 200; time values use 100 × seconds ÷
(seconds + 30), with direction applied afterward. Adding a model does not change another model's
normalized result. Versions, tools, effort and harness conditions remain separate benchmark IDs or result notes.
Multiple submissions from one source for a benchmark are averaged first; source means are combined with reliability weights. Lab-reported evidence has weight 0.35; inactive sources receive a 0.5 factor. Each pillar averages its reported benchmarks using benchmark weight and evidence reliability. Missing pillars stay unknown; this method does not impute them.
The reported pillar mean is reweighted over available pillars, then shrunk toward a neutral prior of 50. Support is the minimum of expected-source coverage, reported pillar weight, and evidence breadth (weighted benchmark evidence divided by a target of four, capped at one). SI Score = 50 + support × (reported pillar mean − 50). This reduces the influence of sparse evidence; it does not remove all differences in evaluation conditions. A rank requires at least 50% confidence and two scored pillars. The fixture used during development describes an older method; production scores always come from the pipeline.
The confidence %, in plain words
Sources publish at different times: an arena board may list a new model on day one, a benchmark steward may take a week, and some sources never cover some models. So every score on this site carries a confidence percentage that starts low and rises as expected sources report.
Each model has a set of expected sources with weights (for example, a frontier model keeps valid reported benchmark evidence in its confidence numerator and denominator even when a source goes stale or fails. A source with no result is expected only when active and publishing on or after the model’s release; catalog and pricing sources do not contribute to capability confidence). The model's coverage is the share of expected weight that has arrived. Confidence is coverage divided by a completeness threshold of 80%, capped at 100:
confidence = min(100, 100 × coverage ÷ 0.8)
In words: a model shows 100% confidence once about 80% of its expected source weight has reported — we do not wait for the last stragglers, because some sources never cover some models. Example: Qwen2.5-Coder-32B-Instruct has pending sources: Official model cards via models.dev and LiveBench; its coverage is 72%, so it shows 90% confidence. Model pages list exactly which sources are in (with arrival dates) and which are pending, and the status page lists every model waiting on any source.
Catalog scope and price conditions
The default leaderboard and provisional pool cover general text LLMs. The model catalog also includes image, video, audio, speech, embedding and robotics products, with their kind shown and a kind filter. Image input alone does not exclude a text model. Specialist generators remain outside the text ranking even when they also produce auxiliary text.
Prices are exact-model on-demand token rates, with the source and retrieval date attached. Context bracket, deployment region, thinking mode and promotions can change the applicable rate; check the source note before estimating a workload. Alibaba uses Global or International USD and the named output mode, Meta uses Standard rather than its contributor discount, and displayed provider promotions remain labeled. Missing rates stay unknown. An unchanged MIT entry may describe an older provider rate even when our retrieval is fresh.
Dated snapshots without their own results are aliases of an existing base product. Evaluated checkpoints and distinct mini, flash, pro and reasoning products retain their identities. Folding an alias does not transfer its price to the base. Provisional catalog rows explain whether results are absent, pillar coverage is insufficient or completeness falls below the unchanged 50% rank floor.
Sources and licenses
Every displayed value links to the source it came from, with the retrieval date. 5 of our 29 sources publish under an open license; the rest are cited as factual figures (prices, release dates, benchmark results) from the publisher's own page.
- models.dev catalog MIT models.dev; MIT; source links accompany every observation.
- Anthropic Mythos 5 system-card claims lab-reported factual citation Anthropic Mythos 5 system-card claims; factual citation; source links accompany every observation.
- Official model cards via models.dev lab-reported factual citation; MIT transcription Official model cards via models.dev; factual citation; MIT transcription; source links accompany every observation.
- LiteLLM pricing MIT LiteLLM; MIT; source links accompany every observation.
- Epoch AI Benchmarking benchmark CC-BY Epoch AI Benchmarking; CC-BY; source links accompany every observation.
- LMArena / Arena preference CC-BY-4.0 LMArena / Arena; CC-BY-4.0; source links accompany every observation.
- SWE-bench Verified benchmark factual citation SWE-bench Verified; factual citation; source links accompany every observation.
- SWE-bench Pro (public) benchmark factual citation SWE-bench Pro (public); factual citation; source links accompany every observation.
- ARC Prize benchmark factual citation ARC Prize; factual citation; source links accompany every observation.
- Humanity’s Last Exam benchmark factual citation Humanity’s Last Exam; factual citation; source links accompany every observation.
- LiveBench benchmark factual citation; Apache-2.0 code LiveBench; factual citation; Apache-2.0 code; source links accompany every observation.
- Aider polyglot benchmark Apache-2.0 Aider polyglot; Apache-2.0; source links accompany every observation.
- Terminal-Bench benchmark Apache-2.0; factual citation Terminal-Bench; Apache-2.0; factual citation; source links accompany every observation.
- OpenRouter rankings popularity CC-BY-4.0 OpenRouter rankings; CC-BY-4.0; source links accompany every observation.
- Hugging Face Hub catalog factual metadata; model-specific licenses Hugging Face Hub; factual metadata; model-specific licenses; source links accompany every observation.
- Anthropic models & pricing pricing factual citation Anthropic models & pricing; factual citation; source links accompany every observation.
- OpenAI API changelog catalog factual citation OpenAI API changelog; factual citation; source links accompany every observation.
- Google Gemini release notes catalog CC-BY-4.0 factual citation Google Gemini release notes; CC-BY-4.0 factual citation; source links accompany every observation.
- DeepSeek V4.1 Flash announcement catalog factual citation DeepSeek V4.1 Flash announcement; factual citation; source links accompany every observation.
- xAI models & pricing pricing factual citation xAI models & pricing; factual citation; source links accompany every observation.
- Provider pricing pages and model cards pricing, lab-reported factual citation Linked individually next to each value
What we deliberately do not use: Artificial Analysis and llm-stats are consulted only as an internal cross-check of our own normalized values — their terms restrict reuse in a competing product, so none of their numbers are published here. Benchmark datasets and questions themselves are third-party IP and are never republished; only results are cited.
Update policy
New models are detected from catalog diffs and provider announcements, and go live as soon as identity and a primary source are confirmed — marked provisional, with unknown fields shown as “not yet reported” rather than borrowed from a sibling model. Scores update as sources report, and the “as of” date in the header shows when the current snapshot was generated. Prices retain the linked catalog or provider provenance; missing prices are not estimated. Direct provider pages take precedence over exact first-party rates transcribed by the MIT models.dev and LiteLLM catalogs. Reseller quotes and sibling-model prices cannot fill a gap; each transcription keeps its original source and retrieval date.
Corrections
Every figure keeps its source and retrieval date, so errors are traceable. When a source corrects a figure we update it and keep the change in the snapshot history; this launch snapshot does not yet publish per-model change histories. Corrections should identify the model, benchmark variant, source URL and retrieval date so they can be reproduced.
Reusing our data
Our composite — the SI Score, confidence and ranks — is published under CC-BY-4.0 with attribution to SuperIndex — the Superintelligence Leaderboard:
- /data/latest.json — the full composite, machine-readable
- /data/models.csv — composite + key facts per model
- /data/benchmarks.csv — the open-licensed benchmark inputs, with per-row source, date and license
Values from “factual citation” sources (for example SWE-bench or provider pricing pages) are shown on this site with attribution but are not included in the open-inputs CSV, because we do not hold a redistribution license for them. Check the original publisher's terms before reusing those.