OpenAI, released Apr 24, 2026

GPT-5.5 Proprice, context, benchmarks and release details

53% confidence 53 percent, Medium confidence, 3 of 7 expected sources in
63.5
SI Score
#40 of 142 ranked models
Input, per 1M tokens
$30.00models.devFirst-party hosted API; MIT models.dev transcription. Provider documentation: https://platform.openai.com/docs/models. Exact canonical endpoint; lowest short-context Standard USD token tier; cache/batch discounts excluded. Deprecated endpoints excluded; openai/gpt-5.5-pro Retrieved Oct 9, 2026 · MIT
Open source ↗
Output, per 1M tokens
$180.00models.devFirst-party hosted API; MIT models.dev transcription. Provider documentation: https://platform.openai.com/docs/models. Exact canonical endpoint; lowest short-context Standard USD token tier; cache/batch discounts excluded. Deprecated endpoints excluded; openai/gpt-5.5-pro Retrieved Oct 9, 2026 · MIT
Open source ↗
Context window
1.1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
128Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Apr 24, 2026OpenAI API changelogPublished source fact Retrieved Oct 9, 2026 · factual citation
Open source ↗

How the score breaks down

Coding (weight 40 percent) —
Math (weight 15 percent) 73.3
Preference (weight 15 percent) —
Reasoning (weight 30 percent) 85.9

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 06:15 UTC.

Around it on the leaderboard

  1. 38 Kimi K2.6 64.0
  2. 39 Qwen3.5 397B-A17B 64.0
  3. 40 GPT-5.5 Pro 63.5
  4. 41 GPT-5.6 Luna 63.4
  5. 42 Gemma 4 26B A4B IT 63.4

Full leaderboard

Benchmark results

10 benchmarks, 14 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

ARC-AGI-1 (public eval)reasoning 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
98.0 2 settings
  • Setting 1 98%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 2 98.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-1 (public eval)
ARC-AGI-1 (semi-private)reasoning 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
96.5 2 settings
  • Setting 1 96.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 2 95%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-1 (semi-private)
ARC-AGI-2 (public eval)reasoning 90.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
90.4 2 settings
  • Setting 1 90.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 2 90.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (public eval)
ARC-AGI-2 (semi-private)reasoning 84.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
84.6 2 settings
  • Setting 1 84.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 2 84.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] gpt-5-5-pro-2026-04-23-xhighPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (semi-private)
Humanity's Last Examreasoning 43.1%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
43.1
Humanity's Last Exam (with tools)reasoning 57.2%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
57.2
FrontierMath Tier 4math 39.6%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Tier 4Published Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
39.6
FrontierMath Tier 4 (v2)math 78.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.0
FrontierMath Tiers 1–3math 52.4%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] Tier 1-3Published Apr 23, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
52.4
FrontierMath Tiers 1–3 (v2)math 87.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Jun 12, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
87.7

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Nomodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
not yet reported
Input modalities
text, image, pdfmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
43% of expected source weight

Reported (3)

  • ARC PrizeOct 8, 2026
  • Epoch AI BenchmarkingOct 8, 2026
  • Official model cards via models.devOct 8, 2026

Awaiting (4)

  • Humanity’s Last Exam13% of weight
  • LiveBench13% of weight
  • LMArena / Arena25% of weight
  • Terminal-Bench6% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed