Moonshot AI, released Jul 11, 2025
Kimi K2 Instructprice, context, benchmarks and release details
- Input, per 1M tokens
- not yet reported
- Output, per 1M tokens
- not yet reported
- Context window
- 131Kmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Max output
- 32.8Kmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - Released
- Jul 11, 2025models.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗
How this score is built
| Pillar | Score | Weight | Adds |
|---|---|---|---|
| Coding | 38.1 | 100% of 40% | 38.1 |
| Math | Not measured yet, left out | ||
| Preference | Not measured yet, left out | ||
| Reasoning | Not measured yet, left out | ||
| Average of measured pillars | 38.1 | ||
| Evidence check | 9.9 (1.0 of 6suites needed) | ||
| SI Score | 48.0 | ||
Measured suites (2)
Related tasks and editions share one suite weight in each pillar, and one reliability-weighted breadth contribution across the model.
- swe-bench-pro: 0.50 breadth; Coding 27.7 × 0.50. 1 result ID: swe-bench-pro-public
- swe-bench-verified: 0.50 breadth; Coding 48.6 × 0.50. 1 result ID: swe-bench-verified
Missing pillars are left out and the remaining weights are rescaled, so nothing counts as a zero. With fewer than 6 weighted suites, the average is pulled toward 50 until more results arrive. Ranking needs at least three measured pillars. Still to report for this model, and not counted against it: ARC Prize, Epoch AI Benchmarking, Humanity’s Last Exam, Official model cards via models.dev, LiveBench, LMArena / Arena, Terminal-Bench. Full method
Method si-v5-suite-evidence-1, computed Oct 9, 2026, 18:16 UTC.
Around it on the leaderboard
- 1 Claude Opus 5.5 80.9
- 2 Claude Fable 5 79.9
- 3 Qwen3.8 Max 77.4
- 4 GPT-6.1 Sol 77.3
- 5 Claude Opus 5 76.9
Benchmark results
2 benchmarks, 3 resultsEach row shows the best published result. Where a model was tested at several settings, such as reasoning effort, open the row to see each one. Hover or tap a value for its source.
Open source ↗ 27.7
SWE-bench Verifiedcoding 53.4%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] CodeSweep - SWE-agentPublished Aug 4, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗
53.4 2 settings
- Setting 1 53.4%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] CodeSweep - SWE-agentPublished Aug 4, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗ - Setting 2 43.8%SWE-bench VerifiedSWE-bench published model plus agent result; harness retained, not a base model evaluation [variant] mini-SWE-agent; 1.7.0Published Aug 7, 2025
Retrieved Oct 9, 2026 · factual citation
Open source ↗
Normalization uses a fixed 0–100 scale for each unit, independent of other models. Compare evaluation conditions before reading a small gap as decisive. “Lab-reported” marks the provider's own published figure.
Details and sources
- Open weights
- YesHugging Face HubPublic Hub repo with weight files; gating/repo upload date is not release date
Retrieved Oct 9, 2026 · factual metadata; model-specific licenses
Open source ↗ - License
- otherHugging Face HubPublished source fact
Retrieved Oct 9, 2026 · factual metadata; model-specific licenses
Open source ↗ - Input modalities
- textmodels.devPublished source fact
Retrieved Oct 9, 2026 · MIT
Open source ↗ - First seen by SuperIndex
- Oct 9, 2026
- Coverage
- 20% of expected source weight
Reported (2)
- SWE-bench VerifiedOct 9, 2026
- SWE-bench Pro (public)Oct 9, 2026
Awaiting (7)
- ARC Prize10% of weight
- Epoch AI Benchmarking20% of weight
- Humanity’s Last Exam10% of weight
- Official model cards via models.dev4% of weight
- LiveBench10% of weight
- LMArena / Arena20% of weight
- Terminal-Bench5% of weight
Confidence rises as pending sources publish. Some sources never cover some models, so confidence reaches 100% at 80% of expected weight.