OpenAI
33 ranked models in a catalog of 57 text models. GPT-6 Astra leads this lab at #3 overall.
Ranked models
Global ranks · current snapshot| Rank | Model | SI Score | coding | math | reasoning | preference | Confidence | Blended price | Context | Released |
|---|---|---|---|---|---|---|---|---|---|---|
| #3 | GPT-6 Astra | 77.5 | 64.8 | 92.9 | 87.1 | 76.9 | 100% | $20.00 | 1.1M | Sep 3, 2026 |
| #6 | GPT-6.1 Sol | 73.0 | 65.0 | 93.0 | 87.5 | 77.4 | 100% | $4.00 | 1.1M | Sep 29, 2026 |
| #7 | GPT-5.6 Sol | 72.1 | 63.8 | 89.3 | 82.1 | 78.4 | 100% | $8.00 | 1.1M | Jul 9, 2026 |
| #13 | GPT-5.4 | 69.7 | 64.9 | 78.5 | 71.6 | 77.9 | 100% | $5.63 | 1.1M | Mar 5, 2026 |
| #14 | GPT-5.5 | 69.5 | 66.6 | 82.1 | 77.6 | 79.2 | 100% | $11.25 | 1.1M | Apr 24, 2026 |
| #16 | GPT-6 Sol | 68.9 | 61.9 | 91.3 | 78.1 | 72.6 | 100% | $4.00 | 1.1M | Sep 22, 2026 |
| #21 | GPT-5.6 Terra | 68.2 | 59.1 | 85.1 | 76.3 | 77.4 | 100% | $4.50 | 1.1M | Jul 9, 2026 |
| #33 | GPT-5.2 | 65.9 | 60.1 | 75.8 | 67.4 | 74.3 | 100% | $4.81 | 400K | Dec 11, 2025 |
| #40 | GPT-5.5 Pro | 63.5 | — | 73.3 | 85.9 | — | 53% | $67.50 | 1.1M | Apr 24, 2026 |
| #41 | GPT-5.6 Luna | 63.4 | 56.7 | 74.5 | 67.0 | 76.0 | 100% | $0.45 | 1.1M | Jul 9, 2026 |
| #55 | GPT-6 Luna | 61.6 | 54.6 | 79.5 | 64.7 | 72.2 | 100% | $0.20 | 1.1M | Sep 22, 2026 |
| #58 | GPT-5.4 Pro | 61.3 | — | 63.6 | 80.8 | — | 69% | $67.50 | 1.1M | Mar 5, 2026 |
| #69 | GPT-5.5 Instant | 58.3 | — | 25.1 | 82.5 | 74.9 | 64% | — | 400K | May 5, 2026 |
| #75 | GPT-5.4 nano | 57.4 | 55.5 | 66.1 | 54.9 | 70.3 | 100% | $0.46 | 400K | Mar 17, 2026 |
| #76 | GPT-5.4 mini | 56.9 | 54.1 | 59.0 | 56.3 | 74.2 | 100% | $1.69 | 400K | Mar 17, 2026 |
| #83 | GPT-5 | 55.7 | 66.7 | 55.6 | 35.2 | 73.7 | 100% | $3.44 | 400K | Aug 7, 2025 |
| #85 | GPT-5.1 | 55.7 | 67.3 | 69.0 | 32.6 | 75.2 | 85% | $3.44 | 400K | Nov 13, 2025 |
| #88 | o3 | 55.4 | 66.3 | 65.3 | 34.6 | 74.1 | 87% | $3.50 | 200K | Apr 16, 2025 |
| #90 | GPT OSS 120BOpen weights | 54.8 | 28.0 | 88.9 | 75.8 | 69.6 | 79% | — | 131K | Aug 5, 2025 |
| #100 | GPT OSS 20BOpen weights | 52.5 | — | 52.1 | 53.3 | 60.7 | 64% | — | 131K | Aug 5, 2025 |
| #103 | o4-mini | 51.2 | 63.4 | 47.2 | 29.9 | 68.3 | 87% | $1.93 | 200K | Apr 16, 2025 |
| #105 | GPT-5 Mini | 50.6 | 62.4 | 48.5 | 26.3 | 70.4 | 99% | $0.69 | 400K | Aug 7, 2025 |
| #109 | GPT-4o (2024-05-13) | 49.4 | — | 36.1 | 48.9 | 62.3 | 93% | $7.50 | 128K | May 13, 2024 |
| #117 | GPT-5 Pro | 45.8 | — | 37.7 | 42.1 | — | 64% | $41.25 | 400K | Oct 6, 2025 |
| #121 | GPT-4o (2024-11-20) | 42.2 | 23.2 | 35.3 | 47.9 | — | 61% | $4.38 | 128K | Nov 20, 2024 |
| #123 | o3-mini | 41.6 | 48.1 | — | 14.2 | 64.5 | 96% | $1.93 | 200K | Dec 20, 2024 |
| #124 | GPT-4.1 | 41.5 | 47.8 | 43.3 | 10.3 | 71.4 | 100% | $3.50 | 1M | Apr 14, 2025 |
| #129 | GPT-5 Nano | 39.1 | 34.8 | 39.0 | 14.7 | 64.5 | 85% | $0.14 | 400K | Aug 7, 2025 |
| #132 | GPT-4.1 mini | 36.9 | 28.2 | 46.5 | 9.7 | 66.9 | 87% | $0.70 | 1M | Apr 14, 2025 |
| #133 | GPT-4o (2024-08-06) | 36.7 | 17.0 | 22.7 | 49.2 | 60.2 | 100% | $4.38 | 128K | Aug 6, 2024 |
| #139 | GPT-4.1 nano | 32.1 | 8.9 | 56.3 | 5.8 | 60.4 | 82% | $0.17 | 1M | Apr 14, 2025 |
| #140 | GPT-4o | 31.2 | 16.7 | — | 1.5 | — | 59% | $4.38 | 128K | May 13, 2024 |
| #142 | GPT-4o mini | 22.9 | 3.6 | 22.7 | 7.5 | 60.6 | 100% | $0.26 | 128K | Jul 18, 2024 |
Pillars use a 0–100 scale. * marks an imputed neutral prior where the model has no scored results in that pillar; it is not a measured benchmark result. Prices are USD per 1M tokens; see the blend and price comparison.
Scores and releases
Provisional models
24 awaiting rank-eligible evidence- GPT-6 Astra (Fast)
No usable open benchmark results found
- GPT-5.6 Cyber
No usable open benchmark results found
- GPT-5.3 Chat (latest)
Results cover 1 pillar; at least 2 required
- GPT-5.3 Codex
Results cover 1 pillar; at least 2 required
- GPT-5.3 Codex Spark
No usable open benchmark results found
- GPT-5.2 Codex
Evidence completeness 43.0%; at least 50% required
- GPT-5.2 Chat
No usable open benchmark results found
- GPT-5.2 Pro
Evidence completeness 48.0%; at least 50% required
- GPT-5.1 Codex Max
No usable open benchmark results found
- GPT-5.1 Chat
No usable open benchmark results found
- GPT-5.1 Codex
Results cover 1 pillar; at least 2 required
- GPT-5.1 Codex mini
No usable open benchmark results found
- GPT OSS Safeguard 120B
No usable open benchmark results found
- GPT OSS Safeguard 20B
No usable open benchmark results found
- GPT-5-Codex
Results cover 1 pillar; at least 2 required
- GPT-5 Chat (latest)
No usable open benchmark results found
- o3-pro
Evidence completeness 22.0%; at least 50% required
- o1-pro
Results cover 1 pillar; at least 2 required
- o1
No usable open benchmark results found
- o3-deep-research
No usable open benchmark results found
- o4-mini-deep-research
No usable open benchmark results found
- GPT-3.5-turbo
No usable open benchmark results found
- GPT-4
No usable open benchmark results found
- GPT-4 Turbo
No usable open benchmark results found
Latest OpenAI news
All AI news ↗How Oracle turns days of work into minutes with ChatGPT and Codex
LegalOn halves Codex costs while maintaining development speed
Pollo AI turns creative ideas into campaigns with OpenAI
Disrupting AI-enabled “false front” operations
Helping teens learn, plan, and shape the future of AI
Radisson Hotel Group brings hotel discovery into ChatGPT