DeepSeek, released May 28, 2025
DeepSeek R1 0528price, context, benchmarks and release details
- Input, per 1M tokens
- not yet reported
- Output, per 1M tokens
- not yet reported
- Context window
- 164Kmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Max output
- 32.8Kmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Released
- May 28, 2025models.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗
How this score is built
| Pillar | Score | Weight | Adds |
|---|---|---|---|
| Coding | 71.4 | 40% | 28.6 |
| Math | 86.6 | 15% | 13.0 |
| Preference | 75.7 | 15% | 11.4 |
| Reasoning | 33.7 | 30% | 10.1 |
| Average of measured pillars | 63.0 | ||
| Evidence check | -1.1 (5.5 of 6suites needed) | ||
| SI Score | 61.9 | ||
Measured suites (6)
Related tasks and editions share one suite weight in each pillar, and one reliability-weighted breadth contribution across the model.
- aider: 0.50 breadth; Coding 71.4 × 0.50. 1 result ID: aider-polyglot
- aime: 1.00 breadth; Math 66.4 × 0.50. 1 result ID: otis-mock-aime-2024-2025
- arc-agi: 1.00 breadth; Reasoning 12.4 × 1.00. 4 result IDs: arc-agi-v1-public-eval, arc-agi-v1-semi-private, arc-agi-v2-public-eval, arc-agi-v2-semi-private
- gpqa: 1.00 breadth; Reasoning 76.3 × 0.50. 1 result ID: gpqa-diamond
- lmarena-text: 1.00 breadth; Preference 75.7 × 1.00. 1 result ID: lmarena-text
- math-level-5: 1.00 breadth; Math 96.6 × 1.00. 1 result ID: math-level-5
With fewer than 6 weighted suites, the average is pulled toward 50 until more results arrive. Still to report for this model, and not counted against it: Humanity’s Last Exam, Official model cards via models.dev, LiveBench, Terminal-Bench. Full method
Method si-v5-suite-evidence-1, computed Oct 9, 2026, 18:16 UTC.
Around it on the leaderboard
- 57 Qwen3.5 35B-A3B 62.4
- 58 GPT-5.5 Instant 62.0
- 59 DeepSeek R1 0528 61.9
- 60 Gemini 3 Pro Preview 61.9
- 61 Muse Spark 1.2 61.8
Benchmark results
9 benchmarks, 9 resultsEach row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.
Open source ↗ 27.0
Open source ↗ 21.2
Open source ↗ 0.3
Open source ↗ 1.1
Open source ↗ 76.3
Open source ↗ 96.6
Open source ↗ 66.4
Open source ↗ 71.4
Open source ↗ 75.7
Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.
Details and sources
- Open weights
- Yesmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - License
- MITmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Input modalities
- textmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - First seen by SuperIndex
- Oct 9, 2026
- Coverage
- 66% of expected source weight
Reported (4)
- Aider polyglotOct 9, 2026
- ARC PrizeOct 9, 2026
- Epoch AI BenchmarkingOct 9, 2026
- LMArena / ArenaOct 9, 2026
Awaiting (4)
- Humanity’s Last Exam12% of weight
- Official model cards via models.dev4% of weight
- LiveBench12% of weight
- Terminal-Bench6% of weight
Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.