DeepSeek, released May 28, 2025

DeepSeek R1 0528price, context, benchmarks and release details

Provisional: not enough results to rank yet Open weights 82% confidence 82 percent, High confidence, 4 of 8 expected sources in
61.9
SI Score
#59 of 125 ranked models
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
164Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
32.8Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
May 28, 2025models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How this score is built

Coding (weight 40 percent) 71.4
Math (weight 15 percent) 86.6
Preference (weight 15 percent) 75.7
Reasoning (weight 30 percent) 33.7
How DeepSeek R1 0528's SI Score is calculated
PillarScoreWeightAdds
Coding71.440%28.6
Math86.615%13.0
Preference75.715%11.4
Reasoning33.730%10.1
Average of measured pillars63.0
Evidence check-1.1 (5.5 of 6suites needed)
SI Score61.9
Measured suites (6)

Related tasks and editions share one suite weight in each pillar, and one reliability-weighted breadth contribution across the model.

  • aider: 0.50 breadth; Coding 71.4 × 0.50. 1 result ID: aider-polyglot
  • aime: 1.00 breadth; Math 66.4 × 0.50. 1 result ID: otis-mock-aime-2024-2025
  • arc-agi: 1.00 breadth; Reasoning 12.4 × 1.00. 4 result IDs: arc-agi-v1-public-eval, arc-agi-v1-semi-private, arc-agi-v2-public-eval, arc-agi-v2-semi-private
  • gpqa: 1.00 breadth; Reasoning 76.3 × 0.50. 1 result ID: gpqa-diamond
  • lmarena-text: 1.00 breadth; Preference 75.7 × 1.00. 1 result ID: lmarena-text
  • math-level-5: 1.00 breadth; Math 96.6 × 1.00. 1 result ID: math-level-5

With fewer than 6 weighted suites, the average is pulled toward 50 until more results arrive. Still to report for this model, and not counted against it: Humanity’s Last Exam, Official model cards via models.dev, LiveBench, Terminal-Bench. Full method

Method si-v5-suite-evidence-1, computed Oct 9, 2026, 18:16 UTC.

Around it on the leaderboard

  1. 57 Qwen3.5 35B-A3B 62.4
  2. 58 GPT-5.5 Instant 62.0
  3. 59 DeepSeek R1 0528 61.9
  4. 60 Gemini 3 Pro Preview 61.9
  5. 61 Muse Spark 1.2 61.8

Full leaderboard

Benchmark results

9 benchmarks, 9 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

ARC-AGI-1 (public eval)reasoning 27.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
27.0
ARC-AGI-1 (semi-private)reasoning 21.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
21.2
ARC-AGI-2 (public eval)reasoning 0.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.3
ARC-AGI-2 (semi-private)reasoning 1.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] deepseek_r1_0528-openrouterPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
1.1
GPQA Diamondreasoning 76.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 29, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
76.3
MATH Level 5math 96.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 29, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
96.6
OTIS Mock AIME 2024–2025math 66.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 29, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
66.4
Aider Polyglotcoding 71.4%Aider polyglotPublished source fact [variant] Aider polyglot; 225 cases; 2 attemptsPublished Jun 6, 2025 Retrieved Oct 9, 2026 · Apache-2.0
Open source ↗
71.4
LMArena Textpreference 1427.6 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
75.7

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Yesmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
MITmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Input modalities
textmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 9, 2026
Coverage
66% of expected source weight

Reported (4)

  • Aider polyglotOct 9, 2026
  • ARC PrizeOct 9, 2026
  • Epoch AI BenchmarkingOct 9, 2026
  • LMArena / ArenaOct 9, 2026

Awaiting (4)

  • Humanity’s Last Exam12% of weight
  • Official model cards via models.dev4% of weight
  • LiveBench12% of weight
  • Terminal-Bench6% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails