Mistral AI, released Jun 10, 2025

Magistral Smallprice, context, benchmarks and release details

Provisional: not enough results to rank yet Open weights 48% confidence 48 percent, Low confidence, 2 of 7 expected sources in
37.1
SI Score
Not ranked yet
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
131Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
8.2Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Jun 10, 2025models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How the score breaks down

Coding (weight 40 percent) —
Math (weight 15 percent) 30.0
Preference (weight 15 percent) —
Reasoning (weight 30 percent) 9.3

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 06:15 UTC.

Around it on the leaderboard

  1. 1 Claude Fable 5.1 80.2
  2. 2 Claude Opus 5.5 78.4
  3. 3 GPT-6 Astra 77.5
  4. 4 Claude Fable 5 76.8
  5. 5 Claude Opus 5 74.9

Full leaderboard

Benchmark results

6 benchmarks, 6 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

ARC-AGI-1 (public eval)reasoning 8.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
8.6
ARC-AGI-1 (semi-private)reasoning 5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.0
ARC-AGI-2 (public eval)reasoning 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0
ARC-AGI-2 (semi-private)reasoning 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] magistral-small-2506Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.0
GPQA Diamondreasoning 56.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
56.1
OTIS Mock AIME 2024–2025math 30%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
30.0

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Yesmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
Apache 2.0models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Input modalities
textmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
38% of expected source weight

Reported (2)

  • ARC PrizeOct 8, 2026
  • Epoch AI BenchmarkingOct 8, 2026

Awaiting (5)

  • Humanity’s Last Exam13% of weight
  • Official model cards via models.dev4% of weight
  • LiveBench13% of weight
  • LMArena / Arena25% of weight
  • Terminal-Bench6% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed