MMLU-ProBenchmark scores and sources
Published result; benchmark version and evaluation conditions remain in the id and result note.
| Row | Model | Result | Normalized (0–100) | Confidence |
|---|---|---|---|---|
| 1 | Sakana Namazu Sakana AI | 90.3%Official model cards via models.devLab-reported; metric percent score; transcribed by MIT models.dev catalog; not independently evaluated [variant]
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 90.3 | 100% confidence 100 percent, Full |
| 2 | DeepSeek V4 Pro DeepSeek | 87.5%Official model cards via models.devLab-reported; metric EM; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 87.5 | 85% confidence 85 percent, High |
| 3 | Nemotron 3 Ultra 550B A55B NVIDIA | 86.8%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Published Jun 4, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 86.8 | 100% confidence 100 percent, Full |
| 4 | DeepSeek V4 Flash DeepSeek | 86.2%Official model cards via models.devLab-reported; metric EM; transcribed by MIT models.dev catalog; not independently evaluated [variant] preview checkpoint; max effort
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 86.2 | 53% confidence 53 percent, Medium |
| 5 | Nemotron 3.5 Lightning 30B A3B NVIDIA | 81.9%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] BF16; reasoning; NeMo Gym / NeMo Evaluator SDK
Retrieved Oct 9, 2026 · factual citation; MIT transcription Open source ↗ | 81.9 | 100% confidence 100 percent, Full |
Results are as published by the source behind each value (hover or tap the number). Benchmark scores are shown individually for every source/variant (a model may have multiple rows), and vary by version, harness and date; the normalized column uses the method’s fixed 0–100 scales and feeds the reasoning pillar of the SI Score.