#1 · Anthropic
80.2SI Score · 100% confidence 100 percent, Full
Highest measured pillars: math (91.5) and reasoning (87.1). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $10.00 / $50.00
- Context window
- 1M
- Weights
- Closed weights
- Evidence
- 50 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources → #2 · Anthropic
78.4SI Score · 100% confidence 100 percent, Full
Highest measured pillars: math (91.5) and reasoning (91.4). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $4.00 / $20.00
- Context window
- 1M
- Weights
- Closed weights
- Evidence
- 45 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources → #3 · OpenAI
77.5SI Score · 100% confidence 100 percent, Full
Highest measured pillars: math (92.9) and reasoning (87.1). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $10.00 / $50.00
- Context window
- 1.1M
- Weights
- Closed weights
- Evidence
- 63 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources → #4 · Anthropic
76.8SI Score · 100% confidence 100 percent, Full
Highest measured pillars: math (91.0) and reasoning (89.4). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $10.00 / $50.00
- Context window
- 1M
- Weights
- Closed weights
- Evidence
- 51 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources → #5 · Anthropic
74.9SI Score · 100% confidence 100 percent, Full
Highest measured pillars: math (86.9) and reasoning (84.3). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $5.00 / $25.00
- Context window
- 1M
- Weights
- Closed weights
- Evidence
- 46 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources → #6 · OpenAI
73.0SI Score · 100% confidence 100 percent, Full
Highest measured pillars: math (93.0) and reasoning (87.5). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $2.00 / $10.00
- Context window
- 1.1M
- Weights
- Closed weights
- Evidence
- 47 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources → #7 · OpenAI
72.1SI Score · 100% confidence 100 percent, Full
Highest measured pillars: math (89.3) and reasoning (82.1). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $4.00 / $20.00
- Context window
- 1.1M
- Weights
- Closed weights
- Evidence
- 56 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources → #8 · Anthropic
71.2SI Score · 100% confidence 100 percent, Full
Highest measured pillars: preference (80.5) and math (78.0). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $5.00 / $25.00
- Context window
- 1M
- Weights
- Closed weights
- Evidence
- 47 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources → #9 · Moonshot AI
71.0SI Score · 100% confidence 100 percent, Full
Highest measured pillars: reasoning (81.3) and preference (79.9). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $3.00 / $15.00
- Context window
- 1M
- Weights
- Open weights
- Evidence
- 38 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources → #10 · Anthropic
70.8SI Score · 100% confidence 100 percent, Full
Highest measured pillars: preference (81.6) and reasoning (78.8). These describe published evaluations, rather than a guarantee on your workload.
- Input / output per 1M tokens
- $5.00 / $25.00
- Context window
- 1M
- Weights
- Closed weights
- Evidence
- 46 published result rows
Prices and context are catalog or provider facts, not measurements by SuperIndex. Missing prices mean a value comparison is unavailable. Open weights do not by themselves establish a permissive license; inspect the model’s license and sources.
Inspect benchmarks and sources →