Claude Fable 5vsQwen3.8 Max

SI Score, benchmarks, price and context compared, with every number sourced.

79.9
SI Score
#2 of 125 ranked models

Anthropic 100% confidence 100 percent, Full confidence

77.4
SI Score
#3 of 125 ranked models

Alibaba / Qwen 85% confidence 85 percent, High confidence

Pillar by pillar

Claude Fable 5Qwen3.8 Max
86.5 Reasoning30% of score 82.6
91.8 Math15% of score 78.7
70.0 Coding40% of score 71.7
81.1 Preference15% of score 80.5

The basics

AttributeClaude Fable 5Qwen3.8 Max
Input price, per 1M tokens $10.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
$1.65Alibaba Model Studio pricingOfficial Alibaba Model Studio Global USD on-demand API; 0<Token≤1M; Non-Thinking and Thinking modes. Output uses non-thinking rate when both modes are available; thinking-only products use their thinking rate. Cache, Batch, free quotas and regional rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Output price, per 1M tokens $50.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
$4.95Alibaba Model Studio pricingOfficial Alibaba Model Studio Global USD on-demand API; 0<Token≤1M; Non-Thinking and Thinking modes. Output uses non-thinking rate when both modes are available; thinking-only products use their thinking rate. Cache, Batch, free quotas and regional rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Context window 1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released Jun 9, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Aug 3, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Weights Closed Closed

Highlighted values are the lower price or the larger context window.

Shared benchmarks

24 in common. Best result ahead: Claude Fable 5 on 18, Qwen3.8 Max on 4

Each side shows its best published result. Results can come from different settings or harnesses, so open a row to compare like with like before reading much into a small gap.

FrontierMath Tier 4 (v2) 90.2% 46.3%

Claude Fable 5

  • max effort90.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 9, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Qwen3.8 Max

  • xhigh effort46.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tier 4 (v2)
FrontierMath Tiers 1–3 (v2) 87.0% 74.7%

Claude Fable 5

  • max effort87.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 9, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Qwen3.8 Max

  • xhigh effort74.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tiers 1–3 (v2)
GPQA Diamond 85.9% 92.7%

Claude Fable 5

  • low effort78.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • high effort83.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort85.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Qwen3.8 Max

  • xhigh effort92.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 292.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
About GPQA Diamond
Humanity's Last Exam 59% 43.6%

Claude Fable 5

  • Published result59%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗

Qwen3.8 Max

  • Published result43.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
About Humanity's Last Exam
Humanity's Last Exam (with tools) 64.5% 56.2%

Claude Fable 5

  • Published result64.5%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗

Qwen3.8 Max

  • Published result56.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
About Humanity's Last Exam (with tools)
LiveBench Coding: code completion 80.4% 73.9%

Claude Fable 5

  • Published result80.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result73.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code completion
LiveBench Coding: code generation 91.5% 71.8%

Claude Fable 5

  • Published result91.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result71.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code generation
LiveBench Coding: JavaScript 68.2% 77.3%

Claude Fable 5

  • Published result68.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result77.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: JavaScript
LiveBench Coding: Python 65% 60%

Claude Fable 5

  • Published result65%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result60%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: Python
LiveBench Coding: TypeScript 53.3% 56.7%

Claude Fable 5

  • Published result53.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result56.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: TypeScript
LiveBench Math: AMPS Hard 99% 98%

Claude Fable 5

  • Published result99%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: AMPS Hard
LiveBench Math: competition math 95.1% 95.1%

Claude Fable 5

  • Published result95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: competition math
LiveBench Math: integrals 97% 81%

Claude Fable 5

  • Published result97%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result81%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: integrals
LiveBench Math: olympiad 92.8% 91.2%

Claude Fable 5

  • Published result92.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result91.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: olympiad
LiveBench Math: simplify 72.0% 67.3%

Claude Fable 5

  • Published result72.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result67.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: simplify
LiveBench Reasoning: connections 99.3% 94.5%

Claude Fable 5

  • Published result99.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result94.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: connections
LiveBench Reasoning: consecutive events 91.4% 87.1%

Claude Fable 5

  • Published result91.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result87.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: consecutive events
LiveBench Reasoning: logic with navigation 78% 74%

Claude Fable 5

  • Published result78%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result74%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: logic with navigation
LiveBench Reasoning: spatial 96% 100%

Claude Fable 5

  • Published result96%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: spatial
LiveBench Reasoning: theory of mind 84.6% 78.8%

Claude Fable 5

  • Published result84.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result78.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: theory of mind
LiveBench Reasoning: zebra puzzles 100% 100%

Claude Fable 5

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Qwen3.8 Max

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: zebra puzzles
LMArena Text 1491.2 elo 1483.3 elo

Claude Fable 5

  • Published result1491.2 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗

Qwen3.8 Max

  • Published result1483.3 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗
About LMArena Text
OTIS Mock AIME 2024–2025 100% 99.4%

Claude Fable 5

  • low effort97.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • high effort100%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort99.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jun 10, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Qwen3.8 Max

  • xhigh effort99.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About OTIS Mock AIME 2024–2025
Terminal-Bench 2.1 88% 86.6%

Claude Fable 5

  • Published result88%Official model cards via models.devLab-reported; metric success rate; transcribed by MIT models.dev catalog; not independently evaluated [variant] 2.1Published Jun 9, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗

Qwen3.8 Max

  • Published result86.6%Official model cards via models.devLab-reported; metric avg@10; transcribed by MIT models.dev catalog; not independently evaluated [variant] 5h timeout; Claude Code; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
About Terminal-Bench 2.1

Which should you choose?

  • For reasoning, Claude Fable 5 leads by 3.8 points.
  • For math, Claude Fable 5 leads by 13.0 points.
  • Qwen3.8 Max costs less per input token ($1.65 vs $10.00 per 1M).
  • Confidence is 100% for Claude Fable 5 and 85% for Qwen3.8 Max; sources still to report can move either score.

These follow from the numbers above. They're not a verdict on your use case.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails