Anthropic, released Feb 19, 2025

Claude Sonnet 3.7price, context, benchmarks and release details

Provisional: not enough results to rank yet 100% confidence 100 percent, Full confidence, 6 of 9 expected sources in
48.9
SI Score
#110 of 142 ranked models
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
200Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
64Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Feb 19, 2025models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How the score breaks down

Coding (weight 40 percent) 60.7
Math (weight 15 percent) 70.9
Preference (weight 15 percent) 62.2
Reasoning (weight 30 percent) 14.4

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 05:13 UTC.

Around it on the leaderboard

  1. 108 Qwen3 Max 50.2
  2. 109 GPT-4o (2024-05-13) 49.4
  3. 110 Claude Sonnet 3.7 48.9
  4. 111 Mistral Small 3.1 24B 48.4
  5. 112 QwQ 32B 48.2

Full leaderboard

Benchmark results

10 benchmarks, 31 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

ARC-AGI-1 (semi-private)reasoning 28.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
28.6 4 settings
  • Setting 1 13.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 2 28.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 3 11.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 1KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 4 21.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 8KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-1 (semi-private)
ARC-AGI-2 (public eval)reasoning 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.8 4 settings
  • Setting 1 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 2 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 3 0.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 1KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 4 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 8KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (public eval)
ARC-AGI-2 (semi-private)reasoning 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 8KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
0.9 4 settings
  • Setting 1 0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7Published Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 2 0.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 16KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 3 0.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 1KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 4 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] Claude 3.7 Thinking 8KPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (semi-private)
GPQA Diamondreasoning 78.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished May 26, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.5 4 settings
  • 16K thinking 76.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Feb 26, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • 32K thinking 76.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Mar 10, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • 64K thinking 78.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished May 26, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 4 66.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 24, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About GPQA Diamond
Humanity's Last Exam (Scale AI)reasoning 8.0%Humanity’s Last ExamThinking budget: 16,000 tokens. Potential contamination warning: This model was evaluated after the public release of HLE, allowing model builder access to the prompts and solutions. [variant] Published Apr 10, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
8.0
MATH Level 5math 91.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Mar 13, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
91.2 4 settings
  • 16K thinking 86.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Feb 26, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • 32K thinking 90.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Mar 12, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • 64K thinking 91.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Mar 13, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 4 68.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 24, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About MATH Level 5
OTIS Mock AIME 2024–2025math 57.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Mar 13, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
57.8 4 settings
  • 16K thinking 46.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished Feb 26, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • 32K thinking 53.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Mar 12, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • 64K thinking 57.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 64KPublished Mar 13, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 4 21.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 25, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About OTIS Mock AIME 2024–2025
Aider Polyglotcoding 60.4%Aider polyglotPublished source fact [variant] Aider polyglot; 225 cases; 2 attemptsPublished Feb 24, 2025 Retrieved Oct 9, 2026 · Apache-2.0
Open source ↗
60.4
SWE-bench Verifiedcoding 66.4%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] Aime-coder v1Published May 14, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
66.4 4 settings
  • Setting 1 61.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Feb 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 2 66.4%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] Aime-coder v1Published May 14, 2025 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 3 52.8%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] mini-SWE-agent; 0.0.0Published Jul 20, 2025 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 4 63.2%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] ToolsPublished Feb 24, 2025 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About SWE-bench Verified
LMArena Textpreference 1299.4 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
62.2

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Nomodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
not yet reported
Input modalities
text, image, pdfmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
80% of expected source weight

Reported (6)

  • Aider polyglotOct 8, 2026
  • ARC PrizeOct 8, 2026
  • Epoch AI BenchmarkingOct 8, 2026
  • Humanity’s Last ExamOct 8, 2026
  • LMArena / ArenaOct 8, 2026
  • SWE-bench VerifiedOct 8, 2026

Awaiting (3)

  • Official model cards via models.dev4% of weight
  • LiveBench11% of weight
  • Terminal-Bench5% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed