OpenAI, released Jan 14, 2026

GPT-5.2 Codexprice, context, benchmarks and release details

Provisional: not enough results to rank yet 43% confidence 43 percent, Low confidence, 4 of 9 expected sources in
57.8
SI Score
Not ranked yet
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
400Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
128Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Jan 14, 2026OpenAI API changelogPublished source fact Retrieved Oct 9, 2026 · factual citation
Open source ↗

How the score breaks down

Coding (weight 40 percent) 62.1
Math (weight 15 percent) 83.1
Preference (weight 15 percent) —
Reasoning (weight 30 percent) 82.5

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 06:15 UTC.

Around it on the leaderboard

  1. 1 Claude Fable 5.1 80.2
  2. 2 Claude Opus 5.5 78.4
  3. 3 GPT-6 Astra 77.5
  4. 4 Claude Fable 5 76.8
  5. 5 Claude Opus 5 74.9

Full leaderboard

Benchmark results

18 benchmarks, 19 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

LiveBench Reasoning: connectionsreasoning 95%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
95.0
LiveBench Reasoning: consecutive eventsreasoning 89.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
89.0
LiveBench Reasoning: logic with navigationreasoning 68%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
68.0
LiveBench Reasoning: spatialreasoning 94%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
94.0
LiveBench Reasoning: theory of mindreasoning 78.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
78.8
LiveBench Reasoning: zebra puzzlesreasoning 70%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
70.0
LiveBench Math: AMPS Hardmath 98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
98.0
LiveBench Math: competition mathmath 95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
95.1
LiveBench Math: integralsmath 75%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
75.0
LiveBench Math: olympiadmath 87.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
87.0
LiveBench Math: simplifymath 60.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
60.6
LiveBench Coding: code completioncoding 87.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
87.0
LiveBench Coding: code generationcoding 80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
80.3
LiveBench Coding: JavaScriptcoding 68.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
68.2
LiveBench Coding: Pythoncoding 60%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
60.0
LiveBench Coding: TypeScriptcoding 20%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
20.0
SWE-bench Pro (public)coding 41.0%Official model cards via models.devLab-reported; metric resolve rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] public Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
41.0 2 settings
  • Setting 1 41.0%Official model cards via models.devLab-reported; metric resolve rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] public Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
  • Setting 2 41.0%SWE-bench Pro (public)Published steward score [variant] Published Jan 27, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About SWE-bench Pro (public)
SWE-bench Verifiedcoding 72.8%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] mini-SWE-agent; 2.0.0Published Feb 19, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
72.8

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Nomodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
not yet reported
Input modalities
text, image, pdfmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
34% of expected source weight

Reported (4)

  • Official model cards via models.devOct 8, 2026
  • LiveBenchOct 8, 2026
  • SWE-bench VerifiedOct 8, 2026
  • SWE-bench Pro (public)Oct 8, 2026

Awaiting (5)

  • ARC Prize10% of weight
  • Epoch AI Benchmarking20% of weight
  • Humanity’s Last Exam10% of weight
  • LMArena / Arena20% of weight
  • Terminal-Bench5% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed