DeepSeek, released May 29, 2025
DeepSeek R1 0528 Qwen3 8Bprice, context, benchmarks and release details
- Input, per 1M tokens
- not yet reported
- Output, per 1M tokens
- not yet reported
- Context window
- 131Kmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Max output
- 32Kmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Released
- May 29, 2025models.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗
How this score is built
| Pillar | Score | Weight | Adds |
|---|---|---|---|
| Math | 43.9 | 33% of 15% | 14.6 |
| Reasoning | 9.3 | 67% of 30% | 6.2 |
| Coding | Not measured yet, left out | ||
| Preference | Not measured yet, left out | ||
| Average of measured pillars | 20.8 | ||
| Evidence check | 19.5 (2.0 of 6suites needed) | ||
| SI Score | 40.3 | ||
Measured suites (2)
Related tasks and editions share one suite weight in each pillar, and one reliability-weighted breadth contribution across the model.
- aime: 1.00 breadth; Math 43.9 × 0.50. 1 result ID: otis-mock-aime-2024-2025
- gpqa: 1.00 breadth; Reasoning 9.3 × 0.50. 1 result ID: gpqa-diamond
Missing pillars are left out and the remaining weights are rescaled, so nothing counts as a zero. With fewer than 6 weighted suites, the average is pulled toward 50 until more results arrive. Ranking needs at least three measured pillars. Still to report for this model, and not counted against it: ARC Prize, Humanity’s Last Exam, Official model cards via models.dev, LiveBench, LMArena / Arena, Terminal-Bench. Full method
Method si-v5-suite-evidence-1, computed Oct 9, 2026, 18:16 UTC.
Around it on the leaderboard
- 1 Claude Opus 5.5 80.9
- 2 Claude Fable 5 79.9
- 3 Qwen3.8 Max 77.4
- 4 GPT-6.1 Sol 77.3
- 5 Claude Opus 5 76.9
Benchmark results
2 benchmarks, 2 resultsEach row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.
Open source ↗ 9.3
Open source ↗ 43.9
Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.
Details and sources
- Open weights
- Yesmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - License
- MITmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Input modalities
- textmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - First seen by SuperIndex
- Oct 9, 2026
- Coverage
- 25% of expected source weight
Reported (1)
- Epoch AI BenchmarkingOct 9, 2026
Awaiting (6)
- ARC Prize13% of weight
- Humanity’s Last Exam13% of weight
- Official model cards via models.dev4% of weight
- LiveBench13% of weight
- LMArena / Arena25% of weight
- Terminal-Bench6% of weight
Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.