Top SI ModelsWhat SI and the SI Score Mean

Understand the initials, the scoring scale and the evidence threshold before using the leading SI models as a shortlist.

Snapshot Oct 9, 2026, 06:15 UTCMethod si-v3-retained-evidence-2

SI is a label; the score is a defined method

On this site SI stands for superintelligence, and SI Score names our composite of published model evaluations. The initials are not a universally agreed scientific unit. A model with an SI Score of 80 is not “80 percent superintelligent,” and doubling a score does not imply twice the intelligence. The number is a convenient synthesis of several measured capabilities under explicit assumptions. It is meaningful only with its method version, source coverage and date.

The term also appears in U.S. policy. Executive Order 14434, signed September 29, 2026, directs executive-branch use of “Super Intelligence” and “SI” for technologies covered by the existing statutory AI definition. That administrative usage differs from the research idea of broadly exceeding human cognitive capabilities. Our model ranking evaluates available systems and does not decide whether the research concept has been achieved.

The current leading SI models

The cards below show the five highest eligible SI Scores in the current snapshot. Their position is determined by the pipeline, not by editorial preference. The full ranked table and provisional catalog are available on the homepage. For ten profiles including price and context, use top super intelligence models.

#1 · Anthropic

Claude Fable 5.1

80.2SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (91.5) and reasoning (87.1). These describe published evaluations, rather than a guarantee on your workload.

Inspect benchmarks and sources →

#2 · Anthropic

Claude Opus 5.5

78.4SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (91.5) and reasoning (91.4). These describe published evaluations, rather than a guarantee on your workload.

Inspect benchmarks and sources →

#3 · OpenAI

GPT-6 Astra

77.5SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (92.9) and reasoning (87.1). These describe published evaluations, rather than a guarantee on your workload.

Inspect benchmarks and sources →

#4 · Anthropic

Claude Fable 5

76.8SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (91.0) and reasoning (89.4). These describe published evaluations, rather than a guarantee on your workload.

Inspect benchmarks and sources →

#5 · Anthropic

Claude Opus 5

74.9SI Score · 100% confidence 100 percent, Full

Highest measured pillars: math (86.9) and reasoning (84.3). These describe published evaluations, rather than a guarantee on your workload.

Inspect benchmarks and sources →

Four pillars, one composite

The snapshot assigns 30% to reasoning, 15% to math, 40% to coding, 15% to preference. Percentage benchmarks retain their absolute scale; unbounded units use fixed transforms. Results from a source are aggregated within each benchmark, then the pillars combine evidence using benchmark and reliability weights. Keeping the scales fixed means adding a new reporter does not automatically change another model’s normalized result.

The method version above shrinks partial evidence toward a neutral prior of 50. Support depends on source coverage, pillar coverage and benchmark breadth. A narrowly tested model therefore cannot turn a single spectacular result into an unqualified headline. Unknown pillars remain unknown; the method does not fabricate a result to fill a gap. The methodology page gives the equations and the source list.

Confidence is completeness

A rank requires at least 50% confidence and results in two pillars. Confidence reaches 100% once the model has received 80% of its expected source weight. This is an operational measure of available evidence, not a statistical confidence interval or probability that the model is correct. Expected sources depend on the pipeline’s coverage and source-activity rules, so compare individual results before drawing a conclusion from completeness alone.

When citing this ranking, include SuperIndex, the method version and the snapshot date. Small gaps can reflect different evaluation conditions, and the ordering can change when a source reports new evidence. Choose the relevant task board for a narrower comparison, and use the published dataset if you need exact values rather than rounded display numbers.

Explore SuperIndex

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed