Claude Opus 5vsClaude Sonnet 5.5

SI Score, benchmarks, price and context compared, with every number sourced.

79.6
SI Score
#4 of 124 ranked models

Anthropic 100% confidence 100 percent, Full confidence

79.5
SI Score
#5 of 124 ranked models

Anthropic 93% confidence 93 percent, High confidence

Pillar by pillar

Claude Opus 5Claude Sonnet 5.5
83.2 Reasoning25% of score 86.4
86.7 Math25% of score 90.6
66.7 Coding25% of score 61.6
82.0 Preference25% of score 79.5

The basics

AttributeClaude Opus 5Claude Sonnet 5.5
Input price, per 1M tokens $5.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
$2.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Output price, per 1M tokens $25.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
$10.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Context window 1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
1MAnthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Released Jul 24, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Sep 28, 2026Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Weights Closed Closed

Highlighted values are the lower price or the larger context window.

Shared benchmarks

23 in common. Best result ahead: Claude Opus 5 on 10, Claude Sonnet 5.5 on 12

Each side shows its best published result. Results can come from different settings or harnesses, so open a row to compare like with like before reading much into a small gap.

FrontierMath Tier 4 (v2) 73.2% 80.5%

Claude Opus 5

  • max effort73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Sonnet 5.5

  • max effort80.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tier 4 (v2)
FrontierMath Tiers 1–3 (v2) 85.6% 88.8%

Claude Opus 5

  • max effort85.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Sonnet 5.5

  • max effort88.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tiers 1–3 (v2)
GPQA Diamond 93.9% 95.6%

Claude Opus 5

  • low effort87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort93.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 392.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Sonnet 5.5

  • max effort95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About GPQA Diamond
Humanity's Last Exam (with tools) 64.7% 64.5%

Claude Opus 5

  • Published result64.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jul 24, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗

Claude Sonnet 5.5

  • Published result64.5%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Sep 28, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
About Humanity's Last Exam (with tools)
LiveBench Coding: code completion 82.6% 84.8%

Claude Opus 5

  • Published result82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result84.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code completion
LiveBench Coding: code generation 80.3% 93.0%

Claude Opus 5

  • Published result80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result93.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code generation
LiveBench Coding: JavaScript 77.3% 36.4%

Claude Opus 5

  • Published result77.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result36.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: JavaScript
LiveBench Coding: Python 75% 35%

Claude Opus 5

  • Published result75%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result35%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: Python
LiveBench Coding: TypeScript 43.3% 46.7%

Claude Opus 5

  • Published result43.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result46.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: TypeScript
LiveBench Math: AMPS Hard 99.0% 98%

Claude Opus 5

  • Published result99.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: AMPS Hard
LiveBench Math: competition math 94.1% 97.1%

Claude Opus 5

  • Published result94.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result97.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: competition math
LiveBench Math: integrals 97% 100%

Claude Opus 5

  • Published result97%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: integrals
LiveBench Math: olympiad 92.8% 91.7%

Claude Opus 5

  • Published result92.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result91.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: olympiad
LiveBench Math: simplify 61.6% 72.4%

Claude Opus 5

  • Published result61.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result72.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: simplify
LiveBench Reasoning: connections 99.3% 98.5%

Claude Opus 5

  • Published result99.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result98.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: connections
LiveBench Reasoning: consecutive events 77.6% 90.9%

Claude Opus 5

  • Published result77.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result90.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: consecutive events
LiveBench Reasoning: logic with navigation 86% 80%

Claude Opus 5

  • Published result86%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result80%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: logic with navigation
LiveBench Reasoning: spatial 100% 98%

Claude Opus 5

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: spatial
LiveBench Reasoning: theory of mind 78.8% 69.2%

Claude Opus 5

  • Published result78.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result69.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: theory of mind
LiveBench Reasoning: zebra puzzles 100% 100%

Claude Opus 5

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: zebra puzzles
LMArena Text 1502.6 elo 1471.2 elo

Claude Opus 5

  • Published result1502.6 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗

Claude Sonnet 5.5

  • Published result1471.2 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗
About LMArena Text
OTIS Mock AIME 2024–2025 98.9% 100%

Claude Opus 5

  • low effort93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort98.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 397.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Sonnet 5.5

  • max effort100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About OTIS Mock AIME 2024–2025
Terminal-Bench 4.0 53.9% 70.6%

Claude Opus 5

  • Setting 153.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; xhighPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
    Open source ↗
  • Setting 251.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Sep 3, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
    Open source ↗
  • Setting 350.3%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; highPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
    Open source ↗
  • Setting 444.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; mediumPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
    Open source ↗
  • Setting 534.9%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; lowPublished Sep 17, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
    Open source ↗

Claude Sonnet 5.5

  • Setting 170.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Sep 28, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
  • Setting 261.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
    Open source ↗
About Terminal-Bench 4.0

Which should you choose?

  • For reasoning, Claude Sonnet 5.5 leads by 3.2 points.
  • For math, Claude Sonnet 5.5 leads by 3.9 points.
  • For coding, Claude Opus 5 leads by 5.1 points.
  • Claude Sonnet 5.5 costs less per input token ($2.00 vs $5.00 per 1M).
  • Confidence is 100% for Claude Opus 5 and 93% for Claude Sonnet 5.5; sources still to report can move either score.

These follow from the numbers above. They're not a verdict on your use case.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails