xAI
5 ranked models in a catalog of 10 text models. Grok 4.6 leads this lab at #27 overall.
Ranked models
Global ranks · current snapshot| Rank | Model | SI Score | coding | math | reasoning | preference | Confidence | Blended price | Context | Released |
|---|---|---|---|---|---|---|---|---|---|---|
| #27 | Grok 4.6 | 66.6 | 57.5 | 78.2 | 76.6 | 75.8 | 100% | $3.00 | 500K | Aug 12, 2026 |
| #36 | Grok 4.7 | 65.5 | 59.0 | 76.2 | 72.7 | 73.0 | 100% | $3.00 | 500K | Sep 21, 2026 |
| #37 | Grok 4.5 | 64.8 | 55.3 | 75.2 | 73.1 | 77.6 | 100% | $3.00 | 500K | Jul 8, 2026 |
| #73 | Grok 4.20 (Reasoning) | 57.7 | — | 43.2 | 65.2 | 77.8 | 80% | $1.56 | 1M | Mar 9, 2026 |
| #92 | Grok 4.3 | 54.6 | 39.1 | 66.5 | 69.2 | 72.8 | 80% | $1.56 | 1M | Apr 17, 2026 |
Pillars use a 0–100 scale. * marks an imputed neutral prior where the model has no scored results in that pillar; it is not a measured benchmark result. Prices are USD per 1M tokens; see the blend and price comparison.
Scores and releases
Provisional models
5 awaiting rank-eligible evidence- Grok Voice STT 1.0
No usable open benchmark results found
- Grok Build 0.1
Evidence completeness 16.0%; at least 50% required
- Grok 4.20 (Non-Reasoning)
No usable open benchmark results found
- Grok 4.1 Fast
No usable open benchmark results found
- Grok 4.1 Fast (Reasoning)
Results cover 1 pillar; at least 2 required