Qwen3.8 MaxvsClaude Opus 5

SI Score, benchmarks, price and context compared, with every number sourced.

77.4
SI Score
#3 of 125 ranked models

Alibaba / Qwen 85% confidence 85 percent, High confidence

76.9
SI Score
#5 of 125 ranked models

Anthropic 100% confidence 100 percent, Full confidence

Pillar by pillar

Qwen3.8 MaxClaude Opus 5
82.6 Reasoning30% of score 83.2
78.7 Math15% of score 86.7
71.7 Coding40% of score 66.7
80.5 Preference15% of score 82.0

The basics

AttributeQwen3.8 MaxClaude Opus 5
Input price, per 1M tokens $1.65Alibaba Model Studio pricingOfficial Alibaba Model Studio Global USD on-demand API; 0<Token≤1M; Non-Thinking and Thinking modes. Output uses non-thinking rate when both modes are available; thinking-only products use their thinking rate. Cache, Batch, free quotas and regional rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
$5.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Output price, per 1M tokens $4.95Alibaba Model Studio pricingOfficial Alibaba Model Studio Global USD on-demand API; 0<Token≤1M; Non-Thinking and Thinking modes. Output uses non-thinking rate when both modes are available; thinking-only products use their thinking rate. Cache, Batch, free quotas and regional rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
$25.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
Context window 1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released Aug 3, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Jul 24, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Weights Closed Closed

Highlighted values are the lower price or the larger context window.

Shared benchmarks

23 in common. Best result ahead: Qwen3.8 Max on 5, Claude Opus 5 on 14

Each side shows its best published result. Results can come from different settings or harnesses, so open a row to compare like with like before reading much into a small gap.

FrontierMath Tier 4 (v2) 46.3% 73.2%

Qwen3.8 Max

  • xhigh effort46.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Opus 5

  • max effort73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tier 4 (v2)
FrontierMath Tiers 1–3 (v2) 74.7% 85.6%

Qwen3.8 Max

  • xhigh effort74.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Opus 5

  • max effort85.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tiers 1–3 (v2)
GPQA Diamond 92.7% 93.9%

Qwen3.8 Max

  • xhigh effort92.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 292.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗

Claude Opus 5

  • low effort87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort93.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 392.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About GPQA Diamond
Humanity's Last Exam 43.6% 56.3%

Qwen3.8 Max

  • Published result43.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗

Claude Opus 5

  • Published result56.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jul 24, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
About Humanity's Last Exam
Humanity's Last Exam (with tools) 56.2% 64.7%

Qwen3.8 Max

  • Published result56.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗

Claude Opus 5

  • Published result64.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jul 24, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
About Humanity's Last Exam (with tools)
LiveBench Coding: code completion 73.9% 82.6%

Qwen3.8 Max

  • Published result73.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code completion
LiveBench Coding: code generation 71.8% 80.3%

Qwen3.8 Max

  • Published result71.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code generation
LiveBench Coding: JavaScript 77.3% 77.3%

Qwen3.8 Max

  • Published result77.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result77.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: JavaScript
LiveBench Coding: Python 60% 75%

Qwen3.8 Max

  • Published result60%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result75%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: Python
LiveBench Coding: TypeScript 56.7% 43.3%

Qwen3.8 Max

  • Published result56.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result43.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: TypeScript
LiveBench Math: AMPS Hard 98% 99.0%

Qwen3.8 Max

  • Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result99.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: AMPS Hard
LiveBench Math: competition math 95.1% 94.1%

Qwen3.8 Max

  • Published result95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result94.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: competition math
LiveBench Math: integrals 81% 97%

Qwen3.8 Max

  • Published result81%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result97%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: integrals
LiveBench Math: olympiad 91.2% 92.8%

Qwen3.8 Max

  • Published result91.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result92.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: olympiad
LiveBench Math: simplify 67.3% 61.6%

Qwen3.8 Max

  • Published result67.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result61.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: simplify
LiveBench Reasoning: connections 94.5% 99.3%

Qwen3.8 Max

  • Published result94.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result99.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: connections
LiveBench Reasoning: consecutive events 87.1% 77.6%

Qwen3.8 Max

  • Published result87.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result77.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: consecutive events
LiveBench Reasoning: logic with navigation 74% 86%

Qwen3.8 Max

  • Published result74%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result86%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: logic with navigation
LiveBench Reasoning: spatial 100% 100%

Qwen3.8 Max

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: spatial
LiveBench Reasoning: theory of mind 78.8% 78.8%

Qwen3.8 Max

  • Published result78.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result78.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: theory of mind
LiveBench Reasoning: zebra puzzles 100% 100%

Qwen3.8 Max

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Claude Opus 5

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: zebra puzzles
LMArena Text 1483.3 elo 1502.6 elo

Qwen3.8 Max

  • Published result1483.3 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗

Claude Opus 5

  • Published result1502.6 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗
About LMArena Text
OTIS Mock AIME 2024–2025 99.4% 98.9%

Qwen3.8 Max

  • xhigh effort99.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Claude Opus 5

  • low effort93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort98.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 397.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About OTIS Mock AIME 2024–2025

Which should you choose?

  • For math, Claude Opus 5 leads by 7.9 points.
  • For coding, Qwen3.8 Max leads by 5.0 points.
  • Qwen3.8 Max costs less per input token ($1.65 vs $5.00 per 1M).
  • Confidence is 85% for Qwen3.8 Max and 100% for Claude Opus 5; sources still to report can move either score.

These follow from the numbers above. They're not a verdict on your use case.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails