Moonshot AI, released Jul 11, 2025

Kimi K2 Instructprice, context, benchmarks and release details

Provisional: not enough results to rank yet Open weights 25% confidence 25 percent, Low confidence, 2 of 9 expected sources in
48.0
SI Score
Not ranked yet
Input, per 1M tokens
not yet reported
Output, per 1M tokens
not yet reported
Context window
131Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Max output
32.8Kmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
Released
Jul 11, 2025models.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗

How this score is built

Coding (weight 40 percent) 38.1
Math (weight 15 percent) —
Preference (weight 15 percent) —
Reasoning (weight 30 percent) —
How Kimi K2 Instruct's SI Score is calculated
PillarScoreWeightAdds
Coding38.1100% of 40%38.1
MathNot measured yet, left out
PreferenceNot measured yet, left out
ReasoningNot measured yet, left out
Average of measured pillars38.1
Evidence check9.9 (1.0 of 6suites needed)
SI Score48.0
Measured suites (2)

Related tasks and editions share one suite weight in each pillar, and one reliability-weighted breadth contribution across the model.

  • swe-bench-pro: 0.50 breadth; Coding 27.7 × 0.50. 1 result ID: swe-bench-pro-public
  • swe-bench-verified: 0.50 breadth; Coding 48.6 × 0.50. 1 result ID: swe-bench-verified

Missing pillars are left out and the remaining weights are rescaled, so nothing counts as a zero. With fewer than 6 weighted suites, the average is pulled toward 50 until more results arrive. Ranking needs at least three measured pillars. Still to report for this model, and not counted against it: ARC Prize, Epoch AI Benchmarking, Humanity’s Last Exam, Official model cards via models.dev, LiveBench, LMArena / Arena, Terminal-Bench. Full method

Method si-v5-suite-evidence-1, computed Oct 9, 2026, 18:16 UTC.

Around it on the leaderboard

  1. 1 Claude Opus 5.5 80.9
  2. 2 Claude Fable 5 79.9
  3. 3 Qwen3.8 Max 77.4
  4. 4 GPT-6.1 Sol 77.3
  5. 5 Claude Opus 5 76.9

Full leaderboard

Benchmark results

2 benchmarks, 3 results

Each row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.

SWE-bench Pro (public)coding 27.7%SWE-bench Pro (public)Published steward score [variant] Published Sep 19, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
27.7
SWE-bench Verifiedcoding 53.4%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] CodeSweep - SWE-agentPublished Aug 4, 2025 Retrieved Oct 9, 2026 · factual citation
Open source ↗
53.4 2 settings
  • Setting 1 53.4%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] CodeSweep - SWE-agentPublished Aug 4, 2025 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
  • Setting 2 43.8%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] mini-SWE-agent; 1.7.0Published Aug 7, 2025 Retrieved Oct 9, 2026 · factual citation
    Open source ↗
About SWE-bench Verified

Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.

Details and sources

Open weights
YesHugging Face HubPublic Hub repo with weight files; gating/repo upload date is not release date Retrieved Oct 9, 2026 · factual metadata; model-specific licenses
Open source ↗
License
otherHugging Face HubPublished source fact Retrieved Oct 9, 2026 · factual metadata; model-specific licenses
Open source ↗
Input modalities
textmodels.devPublished source fact Retrieved Oct 9, 2026 · MIT
Open source ↗
First seen by SuperIndex
Oct 9, 2026
Coverage
20% of expected source weight

Reported (2)

  • SWE-bench VerifiedOct 9, 2026
  • SWE-bench Pro (public)Oct 9, 2026

Awaiting (7)

  • ARC Prize10% of weight
  • Epoch AI Benchmarking20% of weight
  • Humanity’s Last Exam10% of weight
  • Official model cards via models.dev4% of weight
  • LiveBench10% of weight
  • LMArena / Arena20% of weight
  • Terminal-Bench5% of weight

Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.

What changed

Releases · RSS feed

Browser alerts

What to be alerted about
RSS feed

Release emails