DeepSeek, released Jul 31, 2026

DeepSeek V4 Flash 0731price, context, benchmarks and release details

Provisional: not enough results to rank yet Open weights 69% confidence 69 percent, Medium confidence, 4 of 7 expected sources in
61.0
SI Score
#60 of 142 ranked models
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
384Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Jul 31, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How the score breaks down

Coding (weight 40 percent) 59.7
Math (weight 15 percent) 71.3
Preference (weight 15 percent) —
Reasoning (weight 30 percent) 82.9

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 05:13 UTC.

Around it on the leaderboard

  1. 58 GPT-5.4 Pro 61.3
  2. 59 Nemotron 3 Ultra 550B A55B 61.2
  3. 60 DeepSeek V4 Flash 0731 61.0
  4. 61 Hy3 60.7
  5. 62 Claude Haiku 4.5 60.5

Full leaderboard

Benchmark results

25 benchmarks, 33 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

ARC-AGI-1 (public eval)reasoning 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.8 3 settings
  • low effort 93%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 92.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • max effort 94.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-1 (public eval)
ARC-AGI-1 (semi-private)reasoning 89%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
89.0 3 settings
  • low effort 84%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 87%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • max effort 89%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-1 (semi-private)
ARC-AGI-2 (public eval)reasoning 63.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
63.9 3 settings
  • low effort 47.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 59.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • max effort 63.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (public eval)
ARC-AGI-2 (semi-private)reasoning 61.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
61.4 3 settings
  • low effort 46.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort. [variant] deepseek-v4-flash-0731-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 56.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort. [variant] deepseek-v4-flash-0731-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • max effort 61.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort. [variant] deepseek-v4-flash-0731-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (semi-private)
GPQA Diamondreasoning 91.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.0
LiveBench Reasoning: connectionsreasoning 97.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
97.3
LiveBench Reasoning: consecutive eventsreasoning 89.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
89.4
LiveBench Reasoning: logic with navigationreasoning 70%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
70.0
LiveBench Reasoning: spatialreasoning 90%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
90.0
LiveBench Reasoning: theory of mindreasoning 86.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
86.5
LiveBench Reasoning: zebra puzzlesreasoning 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
100.0
FrontierMath Tier 4 (v2)math 24.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
24.4
FrontierMath Tiers 1–3 (v2)math 57.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
57.5
LiveBench Math: AMPS Hardmath 97%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
97.0
LiveBench Math: competition mathmath 96.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
96.1
LiveBench Math: integralsmath 65%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
65.0
LiveBench Math: olympiadmath 89.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
89.1
LiveBench Math: simplifymath 58.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
58.4
OTIS Mock AIME 2024–2025math 94.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 2, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
94.4
LiveBench Coding: code completioncoding 73.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
73.9
LiveBench Coding: code generationcoding 76.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
76.1
LiveBench Coding: JavaScriptcoding 63.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
63.6
LiveBench Coding: Pythoncoding 50%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
50.0
LiveBench Coding: TypeScriptcoding 26.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
26.7
Terminal-Bench 2.1coding 82.7%Official model cards via models.devLab-reported; metric pass@1; transcribed by MIT models.dev catalog; not independently evaluated [variant] max; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
82.7

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
YesHugging Face HubPublic Hub repo with weight files; gating/repo upload date is not release date Retrieved Oct 9, 2026 · factual metadata; model-specific licenses
Open source ↗
License
mitHugging Face HubPublished source fact Retrieved Oct 9, 2026 · factual metadata; model-specific licenses
Open source ↗
Input modalities
textmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Popularity
#8 by usageOpenRouter rankingsPublished source fact; tokens processed; page default period; excludes catalog and endpoint statistics Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
55% of expected source weight

Reported (4)

  • ARC PrizeOct 8, 2026
  • Epoch AI BenchmarkingOct 8, 2026
  • Official model cards via models.devOct 8, 2026
  • LiveBenchOct 8, 2026

Awaiting (3)

  • Humanity’s Last Exam13% of weight
  • LMArena / Arena25% of weight
  • Terminal-Bench6% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed