GPT-6.1 SolvsClaude Sonnet 5.5

SI Score, benchmarks, price and context compared, with every number sourced.

80.8
SI Score
#3 of 124 ranked models

OpenAI 100% confidence 100 percent, Full confidence

79.5
SI Score
#5 of 124 ranked models

Anthropic 93% confidence 93 percent, High confidence

Pillar by pillar

GPT-6.1 SolClaude Sonnet 5.5
88.3 Reasoning25% of score 86.4
95.1 Math25% of score 90.6
62.3 Coding25% of score 61.6
77.4 Preference25% of score 79.5

The basics

AttributeGPT-6.1 SolClaude Sonnet 5.5
Input price, per 1M tokens $2.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts Retrieved Oct 9, 2026 · factual citation
Open source ↗
$2.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Output price, per 1M tokens $10.00OpenAI pricingOfficial Standard short-context rate; excludes Batch/Flex/cache discounts Retrieved Oct 9, 2026 · factual citation
Open source ↗
$10.00Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Context window 1.1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
1MAnthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Released Sep 29, 2026OpenAI API changelogPublished source fact Retrieved Oct 9, 2026 · factual citation
Open source ↗
Sep 28, 2026Anthropic models & pricingOfficial Claude API; lowest short-context on-demand tier; batch/cache/long-context rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Weights Closed Closed

Highlighted values are the lower price or the larger context window.

Shared benchmarks

22 in common. Best result ahead: GPT-6.1 Sol on 9, Claude Sonnet 5.5 on 7

Each side shows its best published result. Results can come from different settings or harnesses, so open a row to compare like with like before reading much into a small gap.

FrontierMath Tier 4 (v2) 100% 80.5%

GPT-6.1 Sol

  • max effort100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Sonnet 5.5

  • max effort80.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tier 4 (v2)
FrontierMath Tiers 1–3 (v2) 93.7% 88.8%

GPT-6.1 Sol

  • max effort93.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Sonnet 5.5

  • max effort88.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tiers 1–3 (v2)
GPQA Diamond 95.4% 95.6%

GPT-6.1 Sol

  • max effort95.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Sonnet 5.5

  • max effort95.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About GPQA Diamond
LiveBench Coding: code completion 82.6% 84.8%

GPT-6.1 Sol

  • Published result82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result84.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code completion
LiveBench Coding: code generation 78.9% 93.0%

GPT-6.1 Sol

  • Published result78.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result93.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code generation
LiveBench Coding: JavaScript 63.6% 36.4%

GPT-6.1 Sol

  • Published result63.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result36.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: JavaScript
LiveBench Coding: Python 60% 35%

GPT-6.1 Sol

  • Published result60%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result35%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: Python
LiveBench Coding: TypeScript 46.7% 46.7%

GPT-6.1 Sol

  • Published result46.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result46.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: TypeScript
LiveBench Math: AMPS Hard 98% 98%

GPT-6.1 Sol

  • Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: AMPS Hard
LiveBench Math: competition math 97.1% 97.1%

GPT-6.1 Sol

  • Published result97.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result97.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: competition math
LiveBench Math: integrals 99% 100%

GPT-6.1 Sol

  • Published result99%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: integrals
LiveBench Math: olympiad 91.9% 91.7%

GPT-6.1 Sol

  • Published result91.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result91.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: olympiad
LiveBench Math: simplify 68.2% 72.4%

GPT-6.1 Sol

  • Published result68.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result72.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: simplify
LiveBench Reasoning: connections 100% 98.5%

GPT-6.1 Sol

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result98.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: connections
LiveBench Reasoning: consecutive events 91.3% 90.9%

GPT-6.1 Sol

  • Published result91.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result90.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: consecutive events
LiveBench Reasoning: logic with navigation 82% 80%

GPT-6.1 Sol

  • Published result82%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result80%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: logic with navigation
LiveBench Reasoning: spatial 98% 98%

GPT-6.1 Sol

  • Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: spatial
LiveBench Reasoning: theory of mind 86.5% 69.2%

GPT-6.1 Sol

  • Published result86.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result69.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: theory of mind
LiveBench Reasoning: zebra puzzles 100% 100%

GPT-6.1 Sol

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Sonnet 5.5

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: zebra puzzles
LMArena Text 1446.5 elo 1471.2 elo

GPT-6.1 Sol

  • Published result1446.5 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗

Claude Sonnet 5.5

  • Published result1471.2 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗
About LMArena Text
OTIS Mock AIME 2024–2025 100% 100%

GPT-6.1 Sol

  • max effort100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Sonnet 5.5

  • max effort100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Sep 29, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About OTIS Mock AIME 2024–2025
Terminal-Bench 4.0 58.2% 70.6%

GPT-6.1 Sol

  • Published result58.2%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Codex; maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
    Open source ↗

Claude Sonnet 5.5

  • Setting 170.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] 4.0Published Sep 28, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
  • Setting 261.8%Terminal-BenchTerminal-Bench 4.0; published harness submission, 95% CI retained at source [variant] Claude Code; maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · Apache-2.0; factual citation
    Open source ↗
About Terminal-Bench 4.0

Which should you choose?

  • For math, GPT-6.1 Sol leads by 4.5 points.
  • Confidence is 100% for GPT-6.1 Sol and 93% for Claude Sonnet 5.5; sources still to report can move either score.

These follow from the numbers above. They're not a verdict on your use case.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails