Moonshot AI, released Jul 16, 2026

Kimi K3price, context, benchmarks and release details

Provisional: not enough results to rank yet Open weights 100% confidence 100 percent, Full confidence, 5 of 7 expected sources in
71.0
SI Score
#9 of 142 ranked models
Input, per 1M tokens
$3.00models.devFirst-party hosted API; MIT models.dev transcription. Provider documentation: https://platform.moonshot.ai/docs/api/chat. Exact canonical endpoint; lowest short-context Standard USD token tier; cache/batch discounts excluded. Deprecated endpoints excluded; moonshotai/kimi-k3 Retrieved Oct 9, 2026 · MIT
Open source ↗
Output, per 1M tokens
$15.00models.devFirst-party hosted API; MIT models.dev transcription. Provider documentation: https://platform.moonshot.ai/docs/api/chat. Exact canonical endpoint; lowest short-context Standard USD token tier; cache/batch discounts excluded. Deprecated endpoints excluded; moonshotai/kimi-k3 Retrieved Oct 9, 2026 · MIT
Open source ↗
Context window
1Mmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
131Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Jul 16, 2026models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How the score breaks down

Coding (weight 40 percent) 71.1
Math (weight 15 percent) 74.5
Preference (weight 15 percent) 79.9
Reasoning (weight 30 percent) 81.3

Weights: reasoning 30%, math 15%, coding 40%, preference 15%. Results use fixed 0–100 scales before averaging, and thin evidence is pulled toward 50. Method si-v3-retained-evidence-2, computed Oct 9, 2026, 06:15 UTC.

Around it on the leaderboard

  1. 7 GPT-5.6 Sol 72.1
  2. 8 Claude Opus 4.7 71.2
  3. 9 Kimi K3 71.0
  4. 10 Claude Opus 4.6 70.8
  5. 11 Muse Spark 1.3 70.1

Full leaderboard

Benchmark results

26 benchmarks, 38 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

ARC-AGI-1 (public eval)reasoning 94.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.9 3 settings
  • low effort 75.8%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 92.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • max effort 94.9%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-1 (public eval)
ARC-AGI-1 (semi-private)reasoning 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
94.5 3 settings
  • low effort 65.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 86.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • max effort 94.5%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-1 (semi-private)
ARC-AGI-2 (public eval)reasoning 61.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
61.7 3 settings
  • low effort 16.1%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 53.2%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • max effort 61.7%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (public eval)
ARC-AGI-2 (semi-private)reasoning 60.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
Open source ↗
60.4 3 settings
  • low effort 12.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'low' reasoning effort via Baseten. [variant] moonshot-kimi-k3-lowPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • high effort 55.0%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'high' reasoning effort via Baseten. [variant] moonshot-kimi-k3-highPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • max effort 60.4%ARC PrizeARC Prize steward-published result; exact edition/subset; publication time is HTTP Last-Modified of the aggregate export, not the evaluation date. Evaluation conducted with 'max' reasoning effort via Baseten. [variant] moonshot-kimi-k3-maxPublished Oct 6, 2026 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About ARC-AGI-2 (semi-private)
GPQA Diamondreasoning 93.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 16, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
93.1 3 settings
  • low effort 84.8%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • high effort 91.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort 93.1%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 16, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About GPQA Diamond
LiveBench Reasoning: connectionsreasoning 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
100.0
LiveBench Reasoning: consecutive eventsreasoning 89.8%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
89.8
LiveBench Reasoning: logic with navigationreasoning 80%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
80.0
LiveBench Reasoning: spatialreasoning 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
100.0
LiveBench Reasoning: theory of mindreasoning 82.7%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
82.7
LiveBench Reasoning: zebra puzzlesreasoning 100%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
100.0
FrontierMath Tier 4 (v2)math 39.0%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 17, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
39.0
FrontierMath Tiers 1–3 (v2)math 72.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 17, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
72.2
LiveBench Math: AMPS Hardmath 97%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
97.0
LiveBench Math: competition mathmath 95.1%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
95.1
LiveBench Math: integralsmath 54%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
54.0
LiveBench Math: olympiadmath 91.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
91.6
LiveBench Math: simplifymath 66.4%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
66.4
OTIS Mock AIME 2024–2025math 97.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 16, 2026 Retrieved Oct 9, 2026 · CC-BY
Open source ↗
97.2 3 settings
  • low effort 68.9%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] lowPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • high effort 93.3%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] highPublished Aug 7, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
  • max effort 97.2%Epoch AI BenchmarkingEpoch-owned evaluation mean score; scores only, no benchmark questions [variant] maxPublished Jul 16, 2026 Retrieved Oct 9, 2026 · CC-BY
    Open source ↗
About OTIS Mock AIME 2024–2025
LiveBench Coding: code completioncoding 82.6%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
82.6
LiveBench Coding: code generationcoding 80.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
80.3
LiveBench Coding: JavaScriptcoding 68.2%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
68.2
LiveBench Coding: Pythoncoding 65%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
65.0
LiveBench Coding: TypeScriptcoding 53.3%LiveBenchLiveBench subtask result; dataset edition 2026_06_25; publication time is HTTP Last-Modified of the score CSV, not the dataset edition or individual evaluation date [variant] 2026_06_25Published Oct 7, 2026 Retrieved Oct 9, 2026 · factual citation; Apache-2.0 code
Open source ↗
53.3
Terminal-Bench 2.1coding 88.3%Official model cards via models.devLab-reported; metric accuracy; transcribed by MIT models.dev catalog; not independently evaluated [variant] max effort; Kimi Code; 2.1Published Jul 16, 2026 Retrieved Oct 9, 2026 · factual citation; MIT transcription
Open source ↗
88.3
LMArena Textpreference 1475.5 eloLMArena / ArenaPublished source fact [variant] text / overallPublished Oct 8, 2026 Retrieved Oct 9, 2026 · CC-BY-4.0
Open source ↗
79.9

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
YesHugging Face HubPublic Hub repo with weight files; gating/repo upload date is not release date Retrieved Oct 9, 2026 · factual metadata; model-specific licenses
Open source ↗
License
otherHugging Face HubPublished source fact Retrieved Oct 9, 2026 · factual metadata; model-specific licenses
Open source ↗
Input modalities
text, image, videomodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 8, 2026
Coverage
81% of expected source weight

Reported (5)

  • ARC PrizeOct 8, 2026
  • Epoch AI BenchmarkingOct 8, 2026
  • Official model cards via models.devOct 8, 2026
  • LiveBenchOct 8, 2026
  • LMArena / ArenaOct 8, 2026

Awaiting (2)

  • Humanity’s Last Exam13% of weight
  • Terminal-Bench6% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

All releases · RSS feed

Alerts on this device

What to be alerted about
RSS feed