Anthropic, released May 22, 2025
Claude Sonnet 4price, context, benchmarks and release details
- Input, per 1M tokens
- $3.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Output, per 1M tokens
- $15.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Context window
- 200Kmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Max output
- 64Kmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Released
- May 22, 2025models.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗
How the score breaks down
Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 06:15 UTC.
Around it on the leaderboard
- 113 DeepSeek-R1 47.9
- 114 GLM-4.6 47.8
- 115 Claude Sonnet 4 46.6
- 116 Nemotron 3.5 Lightning 30B A3B 46.1
- 117 GPT-5 Pro 45.8
Benchmark results
12 benchmarks, 39 resultsEach row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.
ARC-AGI-1 (public eval)reasoning 56.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
56.8 4 settings
- Setting 1 33%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 2 56.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 3 31.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 4 48.6%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
ARC-AGI-1 (semi-private)reasoning 40%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
40.0 4 settings
- Setting 1 23.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 2 40%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 3 28.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 4 29.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
ARC-AGI-2 (public eval)reasoning 4.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
4.4 4 settings
- Setting 1 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 2 4.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 3 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 4 2.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
ARC-AGI-2 (semi-private)reasoning 5.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
5.9 4 settings
- Setting 1 1.3%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514Published Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 2 5.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-16k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 3 0.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-1kPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 4 2.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. [variant] claude-sonnet-4-20250514-thinking-8k-bedrockPublished Oct 6, 2026
Retrieved Oct 9, 2026 · factual citation
Open source ↗
GPQA Diamondreasoning 78.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
78.3 4 settings
- 16K thinking 75.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - 32K thinking 78.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - 59K thinking 77.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 59KPublished May 26, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - Setting 4 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
Open source ↗ 5.5
Open source ↗ 84.4
OTIS Mock AIME 2024–2025math 71.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
71.1 4 settings
- 16K thinking 53.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 16KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - 32K thinking 71.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished May 22, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - 59K thinking 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 59KPublished May 23, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - Setting 4 28.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published May 22, 2025
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
Open source ↗ 56.4
Open source ↗ 9.1
SWE-bench Verifiedcoding 76.8%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] EPAM AI/Run Developer AgentPublished Aug 4, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗
76.8 10 settings
- Setting 1 57%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] Artemis Agent v2Published Sep 24, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 2 71.2%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] BloopPublished Jul 10, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 3 76.8%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] EPAM AI/Run Developer AgentPublished Aug 4, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 4 74.8%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] Harness AIPublished Jul 31, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 5 74.6%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] Lingxi-v1.5Published Jul 20, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 6 64.9%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] mini-SWE-agent; 1.0.0Published Jul 26, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 7 70.8%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] Moatless ToolsPublished Jun 11, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 8 70.4%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] OpenHandsPublished May 24, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 9 66.6%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] SWE-agentPublished May 22, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 10 72.4%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] ToolsPublished May 22, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗
Open source ↗ 66.7
Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.
Details and sources
- Open weights
- Nomodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - License
- not yet reported
- Input modalities
- text, image, pdfmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - First seen by SuperIndex
- Oct 8, 2026
- Coverage
- 82% of expected source weight
Reported (7)
- Aider polyglotOct 8, 2026
- ARC PrizeOct 8, 2026
- Epoch AI BenchmarkingOct 8, 2026
- Humanity’s Last ExamOct 8, 2026
- LMArena / ArenaOct 8, 2026
- SWE-bench VerifiedOct 8, 2026
- SWE-bench Pro (public)Oct 8, 2026
Awaiting (3)
- Official model cards via models.dev3% of weight
- LiveBench10% of weight
- Terminal-Bench5% of weight
Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.