Meta, released Sep 25, 2024

Llama 3.2 1B Instructprice, context, benchmarks and release details

Provisional: not enough results to rank yet Open weights 93% confidence 93 percent, High confidence, 2 of 4 expected sources in
35.1
SI Score
#119 of 125 ranked models
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
131Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
32Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Sep 25, 2024models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How this score is built

Coding (weight 40 percent) —
Math (weight 15 percent) 0.6
Preference (weight 15 percent) 32.6
Reasoning (weight 30 percent) 23.9
How Llama 3.2 1B Instruct's SI Score is calculated
PillarScoreWeightAdds
Math0.625% of 15%0.1
Preference32.625% of 15%8.1
Reasoning23.950% of 30%12.0
CodingNot measured yet, left out
Average of measured pillars20.2
Evidence check14.9 (3.0 of 6suites needed)
SI Score35.1
Measured suites (3)

Related tasks and editions share one suite weight in each pillar, and one reliability-weighted breadth contribution across the model.

  • aime: 1.00 breadth; Math 0.6 × 0.50. 1 result ID: otis-mock-aime-2024-2025
  • gpqa: 1.00 breadth; Reasoning 23.9 × 0.50. 1 result ID: gpqa-diamond
  • lmarena-text: 1.00 breadth; Preference 32.6 × 1.00. 1 result ID: lmarena-text

Missing pillars are left out and the remaining weights are rescaled, so nothing counts as a zero. With fewer than 6 weighted suites, the average is pulled toward 50 until more results arrive. Still to report for this model, and not counted against it: Official model cards via models.dev, LiveBench. Full method

Method si-v5-suite-evidence-1, computed Oct 9, 2026, 18:16 UTC.

Around it on the leaderboard

  1. 117 Llama-3.1-8B-Instruct 36.3
  2. 118 GPT-4.1 mini 35.4
  3. 119 Llama 3.2 1B Instruct 35.1
  4. 120 Llama-3.3-70B-Instruct 34.6
  5. 121 GPT-4o (2024-08-06) 34.4

Full leaderboard

Benchmark results

3 benchmarks, 3 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

GPQA Diamondreasoning 23.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
23.9
OTIS Mock AIME 2024–2025math 0.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 30, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
0.6
LMArena Textpreference 1054.6 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
32.6

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Yesmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
Llama 3.2 Community Licensemodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Input modalities
textmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 9, 2026
Coverage
75% of expected source weight

Reported (2)

  • Epoch AI BenchmarkingOct 9, 2026
  • LMArena / ArenaOct 9, 2026

Awaiting (2)

  • Official model cards via models.dev7% of weight
  • LiveBench19% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails