Meta
9 ranked models in a catalog of 13 text models. Muse Spark 1.3 leads this lab at #11 overall.
Ranked models
Global ranks · current snapshot| Rank | Model | SI Score | coding | math | reasoning | preference | Confidence | Blended price | Context | Released |
|---|---|---|---|---|---|---|---|---|---|---|
| #11 | Muse Spark 1.3 | 70.1 | 63.0 | 83.2 | 90.5 | 81.0 | 93% | $2.00 | 1M | Sep 2, 2026 |
| #47 | Muse Spark 1.1 | 62.9 | 65.4 | 82.9 | 85.3 | 80.1 | 61% | $2.00 | 1M | Apr 8, 2026 |
| #52 | Muse Spark 1.2 | 62.3 | 66.7 | 86.7 | 90.4 | 80.4 | 53% | $2.00 | 1M | Aug 5, 2026 |
| #104 | Llama 4 Scout 17B InstructOpen weights | 51.0 | — | 44.1 | 51.8 | 59.8 | 64% | — | 10M | Apr 5, 2025 |
| #118 | Llama-3.1-70B-InstructOpen weights | 45.7 | — | 25.7 | 44.2 | 57.6 | 93% | — | 128K | Jul 23, 2024 |
| #130 | Llama-3.1-8B-InstructOpen weights | 37.7 | — | 15.8 | 27.0 | 48.3 | 93% | — | 128K | Jul 23, 2024 |
| #137 | Llama-3.3-70B-InstructOpen weights | 32.7 | 3.0 | 29.4 | 47.4 | 59.2 | 100% | — | 128K | Dec 6, 2024 |
| #138 | Llama-3.2-1BOpen weights | 32.1 | — | 0.6 | 23.9 | 32.6 | 93% | — | 131K | Sep 25, 2024 |
| #141 | Llama 4 Maverick 17B InstructOpen weights | 30.0 | 10.4 | 55.5 | 9.2 | — | 78% | — | 1M | Apr 5, 2025 |
Pillars use a 0–100 scale. * marks an imputed neutral prior where the model has no scored results in that pillar; it is not a measured benchmark result. Prices are USD per 1M tokens; see the blend and price comparison.
Scores and releases
Provisional models
4 awaiting rank-eligible evidence- Muse Glimmer 30B
Evidence completeness 6.0%; at least 50% required
- Llama-3.2-11B-Vision-Instruct
No usable open benchmark results found
- Llama-3.2-3B
Results cover 1 pillar; at least 2 required
- Llama-Guard-3-8B
No usable open benchmark results found