Anthropic, released Oct 15, 2025

Claude Haiku 4.5price, context, benchmarks and release details

Provisional: not enough results to rank yet 64% confidence 64 percent, Medium confidence, 2 of 7 expected sources in
60.5
SI Score
#62 of 142 ranked models
Input, per 1M tokens
$1.00models.devFirst-party hosted API; MIT models.dev transcription. Provider documentation: https://docs.anthropic.com/en/docs/about-claude/models. Exact canonical endpoint; lowest short-context Standard USD token tier; cache/batch discounts excluded. Deprecated endpoints excluded; anthropic/claude-haiku-4-5-20251001 Retrieved Oct 9, 2026 · MIT
Open source ↗
Output, per 1M tokens
$5.00models.devFirst-party hosted API; MIT models.dev transcription. Provider documentation: https://docs.anthropic.com/en/docs/about-claude/models. Exact canonical endpoint; lowest short-context Standard USD token tier; cache/batch discounts excluded. Deprecated endpoints excluded; anthropic/claude-haiku-4-5-20251001 Retrieved Oct 9, 2026 · MIT
Open source ↗
Context window
200Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
64Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Oct 15, 2025models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How the score breaks down

Coding (weight 40 percent) —
Math (weight 15 percent) 78.2
Preference (weight 15 percent) 72.7
Reasoning (weight 30 percent) 65.8

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 06:15 UTC.

Around it on the leaderboard

  1. 60 DeepSeek V4 Flash 0731 61.0
  2. 61 Hy3 60.7
  3. 62 Claude Haiku 4.5 60.5
  4. 63 Qwen3.5 Flash 59.5
  5. 64 Claude Haiku 5.5 59.4

Full leaderboard

Benchmark results

4 benchmarks, 7 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

GPQA Diamondreasoning 71.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
71.2 2 settings
  • 32K thinking 71.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 2 60.5%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 16, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About GPQA Diamond
MATH Level 5math 96.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
96.4 2 settings
  • 32K thinking 96.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 2 86.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 16, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About MATH Level 5
OTIS Mock AIME 2024–2025math 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
66.7 2 settings
  • 32K thinking 66.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] 32KPublished Oct 22, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 2 35.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Oct 16, 2025 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About OTIS Mock AIME 2024–2025
LMArena Textpreference 1396.4 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
72.7

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
Nomodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
License
not yet reported
Input modalities
text, image, pdfmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
51% of expected source weight

Reported (2)

  • Epoch AI BenchmarkingOct 8, 2026
  • LMArena / ArenaOct 8, 2026

Awaiting (5)

  • ARC Prize13% of weight
  • Humanity’s Last Exam13% of weight
  • Official model cards via models.dev4% of weight
  • LiveBench13% of weight
  • Terminal-Bench6% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed