Qwen3.8 MaxvsClaude Opus 5
SI Score, benchmarks, price and context compared, with every number sourced.
77.4
SI Score
#3 of 125 ranked models
76.9
SI Score
#5 of 125 ranked models
Pillar by pillar
Qwen3.8 MaxClaude Opus 5
82.6 Reasoning30% of score 83.2
78.7 Math15% of score 86.7
71.7 Coding40% of score 66.7
80.5 Preference15% of score 82.0
The basics
| Attribute | Qwen3.8 Max | Claude Opus 5 |
|---|---|---|
| Input price, per 1M tokens | $1.65Alibaba Model Studio pricingOfficial Alibaba Model Studio Global USD on-demand API; 0<Token≤1M; Non-Thinking and Thinking modes. Output uses non-thinking rate when both modes are available; thinking-only products use their thinking rate. Cache, Batch, free quotas and regional rates excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $5.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Output price, per 1M tokens | $4.95Alibaba Model Studio pricingOfficial Alibaba Model Studio Global USD on-demand API; 0<Token≤1M; Non-Thinking and Thinking modes. Output uses non-thinking rate when both modes are available; thinking-only products use their thinking rate. Cache, Batch, free quotas and regional rates excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ | $25.00Anthropic API pricingOfficial Claude API base tokens; lowest short-context global Standard rate; cache, batch, fast and regional premiums excluded
Retrieved Oct 9, 2026 · factual citation Open source ↗ |
| Context window | 1Mmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ | 1Mmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ |
| Released | Aug 3, 2026models.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ | Jul 24, 2026models.devPublished source fact
Retrieved Oct 9, 2026 · MIT Open source ↗ |
| Weights | Closed | Closed |
Highlighted values are the lower price or the larger context window.
Shared benchmarks
23 in common. Best result ahead: Qwen3.8 Max on 5, Claude Opus 5 on 14Each side shows its best published result. Results can come from different settings or harnesses, so open a row to compare like with like before reading much into a small gap.
FrontierMath Tier 4 (v2) 46.3% 73.2%
Qwen3.8 Max
- xhigh effort46.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
Claude Opus 5
- max effort73.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
FrontierMath Tiers 1–3 (v2) 74.7% 85.6%
Qwen3.8 Max
- xhigh effort74.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
Claude Opus 5
- max effort85.6%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
GPQA Diamond 92.7% 93.9%
Qwen3.8 Max
- xhigh effort92.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - Setting 292.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant]
Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
Claude Opus 5
- low effort87.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - max effort93.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - Setting 392.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
Humanity's Last Exam 43.6% 56.3%
Qwen3.8 Max
- Published result43.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] without tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
Claude Opus 5
- Published result56.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] no toolsPublished Jul 24, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
Humanity's Last Exam (with tools) 56.2% 64.7%
Qwen3.8 Max
- Published result56.2%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] with tools
Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
Claude Opus 5
- Published result64.7%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] with toolsPublished Jul 24, 2026
Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
LiveBench Coding: code completion 73.9% 82.6%
Qwen3.8 Max
- Published result73.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Coding: code generation 71.8% 80.3%
Qwen3.8 Max
- Published result71.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Coding: JavaScript 77.3% 77.3%
Qwen3.8 Max
- Published result77.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result77.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Coding: Python 60% 75%
Qwen3.8 Max
- Published result60%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result75%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Coding: TypeScript 56.7% 43.3%
Qwen3.8 Max
- Published result56.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result43.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: AMPS Hard 98% 99.0%
Qwen3.8 Max
- Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result99.0%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: competition math 95.1% 94.1%
Qwen3.8 Max
- Published result95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result94.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: integrals 81% 97%
Qwen3.8 Max
- Published result81%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result97%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: olympiad 91.2% 92.8%
Qwen3.8 Max
- Published result91.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result92.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Math: simplify 67.3% 61.6%
Qwen3.8 Max
- Published result67.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result61.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: connections 94.5% 99.3%
Qwen3.8 Max
- Published result94.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result99.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: consecutive events 87.1% 77.6%
Qwen3.8 Max
- Published result87.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result77.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: logic with navigation 74% 86%
Qwen3.8 Max
- Published result74%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result86%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: spatial 100% 100%
Qwen3.8 Max
- Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: theory of mind 78.8% 78.8%
Qwen3.8 Max
- Published result78.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result78.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LiveBench Reasoning: zebra puzzles 100% 100%
Qwen3.8 Max
- Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
Claude Opus 5
- Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026
Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
LMArena Text 1483.3 elo 1502.6 elo
Qwen3.8 Max
- Published result1483.3 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
Claude Opus 5
- Published result1502.6 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026
Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
OTIS Mock AIME 2024–2025 99.4% 98.9%
Qwen3.8 Max
- xhigh effort99.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
Claude Opus 5
- low effort93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - max effort98.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 24, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗ - Setting 397.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] Published Aug 6, 2026
Retrieved Oct 9, 2026 · CC-BY
Open source ↗
Which should you choose?
- For math, Claude Opus 5 leads by 7.9 points.
- For coding, Qwen3.8 Max leads by 5.0 points.
- Qwen3.8 Max costs less per input token ($1.65 vs $5.00 per 1M).
- Confidence is 85% for Qwen3.8 Max and 100% for Claude Opus 5; sources still to report can move either score.
These follow from the numbers above. They're not a verdict on your use case.