Moonshot AI

Tier A · Strong with clear operating rules.

Kimi K3

Near-S intelligence held in A tier by real-world cost and quota friction.

Canonical model record

Current identity, limits, and pricing

Provider source · checked 2026-08-14 ↗
Status
Current
API model ID
kimi-k3
Context
1.05M
Max output
131K
API price / 1M tokens
$3 input · $15 output

Suggested for this guide

Kimi

Use the same Kimi family tested on this page. Check the current plan and model access before subscribing.

Best for: Claude Code setups and cost-aware model work

Check Kimi plansPartner link. It supports Superbash Learn at no extra cost to you.

Visual prompt runs

Benchmark runs

Open each generated scene, or compare the same prompt across models.

10 runs

Mechanical Watch Simulator

Interactive watch movement testing mechanical legibility, accurate relative motion, and real-time 3D controls.

Superbash commentary

Our take

Back to the full tier list →

Kimi K3 sits between A and S on capability, but lands in A because it is no longer the cheap Chinese-model bargain the team expected. It integrates broadly and remains highly capable, yet the team exhausted both short-window and total monthly plan allowances by mid-August.

Best for

  • Long-horizon coding in Kimi Code
  • Large-repository and tool-using tasks
  • Visual iteration against screenshots
  • Knowledge work with long context

Watch out

  • Monthly and short-window plan limits affected the team’s use
  • Premium inference pricing
  • Use a K3-compatible harness and preserve thinking history
  • Avoid switching to K3 mid-session

Why it is ranked here

  1. Moonshot reports 67.5 on DeepSWE, 88.3 on Terminal-Bench 2.1, 81.2 on FrontierSWE, and 77.8 on Program Bench at maximum reasoning effort; the comparisons are harness-dependent and should not be treated as a universal leaderboard.
  2. K3’s native multimodal loop is designed to inspect screenshots, refine code, and build interactive output—capabilities reflected in its nine completed Superbash visual benchmark runs.
  3. The team’s move from near-S to A is driven by exhausted subscription quotas and cost, not a loss of confidence in raw intelligence.

Evidence and commentary

2026-08-17

Superbash editorial model ranking

Takeaway: Kimi K3 is currently placed in Tier A.

The August 2026 editorial roster places Kimi K3 at rank 4.

Open source →
2026-07-16

Kimi K3: Open Frontier Intelligence

Takeaway: Moonshot’s release table reports competitive coding, agentic, and vision results, including 88.3% on Terminal-Bench 2.1 and 81.2% on FrontierSWE.

These are vendor-reported results at maximum reasoning effort and vary by agent harness, so they establish K3’s frontier capability but are not a like-for-like overall score.

Open source →
2026-07-16

Superbash visual benchmark suite: Kimi K3

Takeaway: K3 completed all nine shared visual benchmark briefs in the Superbash suite.

The outputs provide a practical check of visual implementation quality alongside published benchmark results.

Open source →

Official benchmark profile

How Kimi K3 scores beyond our visual tests.

Moonshot reports Kimi K3 at maximum reasoning effort across coding, research, tool use, document work, and vision. The official table mixes provider harnesses on some agent tasks, so each row keeps its evaluation setting.

Moonshot AI sourceAugust 2026Source report →
Academic reasoning

GPQA Diamond

93.5%
Coding

DeepSWE

67.5%
Agentic coding

Terminal-Bench 2.1

88.3%
Full official benchmark table12 rows with source settings and peer charts
BenchmarkAreaScoreSetting / comparison
GPQA DiamondAcademic reasoning93.5%Maximum reasoning effort, temperature 1.0, top-p 0.95.
GPT-5.6 Sol94.1%
Kimi K393.5%
Claude Fable 592.6%
GLM 5.291.2%
DeepSWECoding67.5%Kimi Code harness on DeepSWE v1.1 tasks.The official leaderboard score under mini-SWE-agent is 67.3%.
GPT-5.6 Sol73.0%
Claude Fable 570.0%
Kimi K367.5%
GPT-5.567.0%
Claude Opus 4.859.0%
Terminal-Bench 2.1Agentic coding88.3%Kimi Code harness; peers use each provider's cited best harness.
GPT-5.6 Sol88.8%
Kimi K388.3%
Claude Fable 588.0%
Claude Opus 4.884.6%
GPT-5.583.4%
SWE-MarathonLong-horizon coding42.0%Claude Code for Kimi and Claude; Codex for GPT-5.6 on the July 9 H20-calibrated task branch.
Kimi K342.0%
Claude Opus 4.840.0%
GPT-5.6 Sol39.0%
Claude Fable 535.0%
BrowseCompAgentic research91.2%Context compaction at 300K tokens; the full 1M-context run without context management scored 90.4%.
Kimi K391.2%
GPT-5.6 Sol90.4%
Claude Fable 588.0%
GPT-5.584.4%
DeepSearchQA (F1)Deep research95.0%Maximum reasoning effort, temperature 1.0, top-p 1.0.
Kimi K395.0%
Claude Fable 594.2%
Claude Opus 4.893.1%
MCPMark-VerifiedTool use94.5%Maximum reasoning effort, temperature 1.0, top-p 1.0.
Kimi K394.5%
GPT-5.6 Sol92.9%
GPT-5.592.9%
Claude Fable 587.4%
AutomationBenchWorkflow automation30.8%600-task public subset with the official benchmark setup.
Kimi K330.8%
GPT-5.6 Sol29.7%
Claude Fable 529.1%
Claude Opus 4.827.2%
SpreadsheetBench 2Spreadsheet work34.8%Claude Code for Kimi and Claude; Codex for GPT models.
Kimi K334.8%
Claude Fable 534.7%
GPT-5.6 Sol32.4%
Claude Opus 4.831.6%
OmniDocBenchDocument understanding91.1%Maximum reasoning effort; average over three multimodal runs.
Kimi K391.1%
Claude Fable 589.8%
GPT-5.589.4%
Claude Opus 4.887.9%
MMMU-ProMultimodal reasoning81.6% / 83.4%Without tools / with Python; original input order; average over three runs.
GPT-5.6 Sol83.0% / 84.6%
Kimi K381.6% / 83.4%
Claude Fable 581.2% / 86.5%
GPT-5.581.2% / 83.2%
CharXiv (RQ)Chart reasoning84.8% / 91.3%Without tools / with Python; average over three runs.
Claude Fable 588.9% / 93.5%
Kimi K384.8% / 91.3%
GPT-5.6 Sol84.6% / 89.1%
GPT-5.584.1% / 89.0%