Qwen3.8 MaxvsKimi K3

SI Score, benchmarks, price and context compared, with every number sourced.

77.4
SI Score
#3 of 125 ranked models

Alibaba / Qwen 85% confidence 85 percent, High confidence

76.8
SI Score
#6 of 125 ranked models

Moonshot AI 100% confidence 100 percent, Full confidence

Pillar by pillar

Qwen3.8 MaxKimi K3
82.6 Reasoning30% of score 80.5
78.7 Math15% of score 71.9
71.7 Coding40% of score 74.7
80.5 Preference15% of score 79.9

The basics

AttributeQwen3.8 MaxKimi K3
Input price, per 1M tokens $1.65Alibaba Model Studio pricingOfficial Alibaba Model Studio Global USD on-demand API; 0<Token≤1M; Non-Thinking and Thinking modes. Output uses non-thinking rate when both modes are available; thinking-only products use their thinking rate. Cache, Batch, free quotas and regional rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
$3.00models.devFirst-party hosted API; MIT models.dev transcription. Provider documentation: https://platform.moonshot.ai/docs/api/chat. Exact canonical endpoint; lowest short-context Standard USD token tier; cache/batch discounts excluded. Deprecated endpoints excluded; moonshotai/kimi-k3 Retrieved Oct 9, 2026 · MIT
Open source ↗
Output price, per 1M tokens $4.95Alibaba Model Studio pricingOfficial Alibaba Model Studio Global USD on-demand API; 0<Token≤1M; Non-Thinking and Thinking modes. Output uses non-thinking rate when both modes are available; thinking-only products use their thinking rate. Cache, Batch, free quotas and regional rates excluded Retrieved Oct 9, 2026 · factual citation
Open source ↗
$15.00models.devFirst-party hosted API; MIT models.dev transcription. Provider documentation: https://platform.moonshot.ai/docs/api/chat. Exact canonical endpoint; lowest short-context Standard USD token tier; cache/batch discounts excluded. Deprecated endpoints excluded; moonshotai/kimi-k3 Retrieved Oct 9, 2026 · MIT
Open source ↗
Context window 1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released Aug 3, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Jul 16, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Weights Closed Open

Highlighted values are the lower price or the larger context window.

Shared benchmarks

22 in common. Best result ahead: Qwen3.8 Max on 9, Kimi K3 on 10

Each side shows its best published result. Results can come from different settings or harnesses, so open a row to compare like with like before reading much into a small gap.

FrontierMath Tier 4 (v2) 46.3% 39.0%

Qwen3.8 Max

  • xhigh effort46.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Kimi K3

  • max effort39.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 17, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tier 4 (v2)
FrontierMath Tiers 1–3 (v2) 74.7% 72.2%

Qwen3.8 Max

  • xhigh effort74.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Kimi K3

  • max effort72.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 17, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About FrontierMath Tiers 1–3 (v2)
GPQA Diamond 92.7% 93.1%

Qwen3.8 Max

  • xhigh effort92.7%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • Setting 292.6%Official model cards via models.devLab-reported; metric score; transcribed by MIT models.dev catalog; not independently evaluated [variant] Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗

Kimi K3

  • low effort84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • high effort91.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort93.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 16, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About GPQA Diamond
LiveBench Coding: code completion 73.9% 82.6%

Qwen3.8 Max

  • Published result73.9%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code completion
LiveBench Coding: code generation 71.8% 80.3%

Qwen3.8 Max

  • Published result71.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: code generation
LiveBench Coding: JavaScript 77.3% 68.2%

Qwen3.8 Max

  • Published result77.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result68.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: JavaScript
LiveBench Coding: Python 60% 65%

Qwen3.8 Max

  • Published result60%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result65%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: Python
LiveBench Coding: TypeScript 56.7% 53.3%

Qwen3.8 Max

  • Published result56.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result53.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Coding: TypeScript
LiveBench Math: AMPS Hard 98% 97%

Qwen3.8 Max

  • Published result98%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result97%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: AMPS Hard
LiveBench Math: competition math 95.1% 95.1%

Qwen3.8 Max

  • Published result95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: competition math
LiveBench Math: integrals 81% 54%

Qwen3.8 Max

  • Published result81%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result54%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: integrals
LiveBench Math: olympiad 91.2% 91.6%

Qwen3.8 Max

  • Published result91.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result91.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: olympiad
LiveBench Math: simplify 67.3% 66.4%

Qwen3.8 Max

  • Published result67.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result66.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Math: simplify
LiveBench Reasoning: connections 94.5% 100%

Qwen3.8 Max

  • Published result94.5%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: connections
LiveBench Reasoning: consecutive events 87.1% 89.8%

Qwen3.8 Max

  • Published result87.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result89.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: consecutive events
LiveBench Reasoning: logic with navigation 74% 80%

Qwen3.8 Max

  • Published result74%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result80%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: logic with navigation
LiveBench Reasoning: spatial 100% 100%

Qwen3.8 Max

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: spatial
LiveBench Reasoning: theory of mind 78.8% 82.7%

Qwen3.8 Max

  • Published result78.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result82.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: theory of mind
LiveBench Reasoning: zebra puzzles 100% 100%

Qwen3.8 Max

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗

Kimi K3

  • Published result100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
    Open source ↗
About LiveBench Reasoning: zebra puzzles
LMArena Text 1483.3 elo 1475.5 elo

Qwen3.8 Max

  • Published result1483.3 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗

Kimi K3

  • Published result1475.5 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
    Open source ↗
About LMArena Text
OTIS Mock AIME 2024–2025 99.4% 97.2%

Qwen3.8 Max

  • xhigh effort99.4%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] xhighPublished Aug 4, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗

Kimi K3

  • low effort68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • high effort93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort97.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 16, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About OTIS Mock AIME 2024–2025
Terminal-Bench 2.1 86.6% 88.3%

Qwen3.8 Max

  • Published result86.6%Official model cards via models.devLab-reported; metric avg@10; transcribed by MIT models.dev catalog; not independently evaluated [variant] 5h timeout; Claude Code; 2.1 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗

Kimi K3

  • max effort88.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Kimi Code; 2.1Published Jul 16, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
    Open source ↗
About Terminal-Bench 2.1

Which should you choose?

  • For math, Qwen3.8 Max leads by 6.8 points.
  • Qwen3.8 Max costs less per input token ($1.65 vs $3.00 per 1M).
  • Kimi K3 has open weights, which matters if you need to self-host or fine-tune.
  • Confidence is 85% for Qwen3.8 Max and 100% for Kimi K3; sources still to report can move either score.

These follow from the numbers above. They're not a verdict on your use case.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails