Microsoft, released Apr 15, 2024

WizardLM 2 8x22Bprice, context, benchmarks and release details

Provisional: not enough results to rank yet Open weights 100% confidence 100 percent, Full confidence, 1 of 1 expected sources in
45.8
SI Score
Not ranked yet
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
65.5Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
8Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Apr 15, 2024models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How this score is built

Coding (weight 40 percent) —
Math (weight 15 percent) 25.7
Preference (weight 15 percent) —
Reasoning (weight 30 percent) 43.4
How WizardLM 2 8x22B's SI Score is calculated
PillarScoreWeightAdds
Math25.733% of 15%8.6
Reasoning43.467% of 30%29.0
CodingNot measured yet, left out
PreferenceNot measured yet, left out
Average of measured pillars37.5
Evidence check8.3 (2.0 of 6suites needed)
SI Score45.8
Measured suites (2)

Related tasks and editions share one suite weight in each pillar, and one reliability-weighted breadth contribution across the model.

  • gpqa: 1.00 breadth; Reasoning 43.4 × 0.50. 1 result ID: gpqa-diamond
  • math-level-5: 1.00 breadth; Math 25.7 × 1.00. 1 result ID: math-level-5

Missing pillars are left out and the remaining weights are rescaled, so nothing counts as a zero. With fewer than 6 weighted suites, the average is pulled toward 50 until more results arrive. Ranking needs at least three measured pillars. Full method

Method si-v5-suite-evidence-1, computed Oct 9, 2026, 18:16 UTC.

Around it on the leaderboard

  1. 1 Claude Opus 5.5 80.9
  2. 2 Claude Fable 5 79.9
  3. 3 Qwen3.8 Max 77.4
  4. 4 GPT-6.1 Sol 77.3
  5. 5 Claude Opus 5 76.9

Full leaderboard

Benchmark results

2 benchmarks, 2 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

GPQA Diamondreasoning 43.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
43.4
MATH Level 5math 25.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Jan 27, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
25.7

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Yesmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
Apache-2.0models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Input modalities
textmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 9, 2026
Coverage
100% of expected source weight

Reported (1)

  • Epoch AI BenchmarkingOct 9, 2026

Awaiting (0)

Every expected source has reported for this model.

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails