DeepSeek, released May 29, 2025

DeepSeek R1 0528 Qwen3 8Bprice, context, benchmarks and release details

Provisional: not enough results to rank yet Open weights 32% confidence 32 percent, Low confidence, 1 of 7 expected sources in
40.3
SI Score
Not ranked yet
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
131Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
32Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
May 29, 2025models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How this score is built

Coding (weight 40 percent) —
Math (weight 15 percent) 43.9
Preference (weight 15 percent) —
Reasoning (weight 30 percent) 9.3
How DeepSeek R1 0528 Qwen3 8B's SI Score is calculated
PillarScoreWeightAdds
Math43.933% of 15%14.6
Reasoning9.367% of 30%6.2
CodingNot measured yet, left out
PreferenceNot measured yet, left out
Average of measured pillars20.8
Evidence check19.5 (2.0 of 6suites needed)
SI Score40.3
Measured suites (2)

Related tasks and editions share one suite weight in each pillar, and one reliability-weighted breadth contribution across the model.

  • aime: 1.00 breadth; Math 43.9 × 0.50. 1 result ID: otis-mock-aime-2024-2025
  • gpqa: 1.00 breadth; Reasoning 9.3 × 0.50. 1 result ID: gpqa-diamond

Missing pillars are left out and the remaining weights are rescaled, so nothing counts as a zero. With fewer than 6 weighted suites, the average is pulled toward 50 until more results arrive. Ranking needs at least three measured pillars. Still to report for this model, and not counted against it: ARC Prize, Humanity’s Last Exam, Official model cards via models.dev, LiveBench, LMArena / Arena, Terminal-Bench. Full method

Method si-v5-suite-evidence-1, computed Oct 9, 2026, 18:16 UTC.

Around it on the leaderboard

  1. 1 Claude Opus 5.5 80.9
  2. 2 Claude Fable 5 79.9
  3. 3 Qwen3.8 Max 77.4
  4. 4 GPT-6.1 Sol 77.3
  5. 5 Claude Opus 5 76.9

Full leaderboard

Benchmark results

2 benchmarks, 2 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

GPQA Diamondreasoning 9.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
9.3
OTIS Mock AIME 2024–2025math 43.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 28, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
43.9

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Yesmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
MITmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Input modalities
textmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 9, 2026
Coverage
25% of expected source weight

Reported (1)

  • Epoch AI BenchmarkingOct 9, 2026

Awaiting (6)

  • ARC Prize13% of weight
  • Humanity’s Last Exam13% of weight
  • Official model cards via models.dev4% of weight
  • LiveBench13% of weight
  • LMArena / Arena25% of weight
  • Terminal-Bench6% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails