MMLU-ProBenchmark scores and sources

Published result; benchmark version and evaluation conditions remain in the id and result note.

pillar: reasoning · weight 0.5 within pillar · unit: % (higher is better) · official board ↗
Row Model Result Normalized (0–100) Confidence
1 Sakana Namazu Sakana AI 90.3%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
90.3 100% confidence 100 percent, Full
2 DeepSeek V4 Pro DeepSeek 87.5%Official model cards via models.devLab-reported; metric EM; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
87.5 85% confidence 85 percent, High
3 Nemotron 3 Ultra 550B A55B NVIDIA 86.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jun 4, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
86.8 100% confidence 100 percent, Full
4 DeepSeek V4 Flash DeepSeek 86.2%Official model cards via models.devLab-reported; metric EM; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
86.2 53% confidence 53 percent, Medium
5 Nemotron 3.5 Lightning 30B A3B NVIDIA 81.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; NeMo Gym / NeMo Evaluator SDK Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
81.9 100% confidence 100 percent, Full

Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed