Moonshot AI Launches Kimi K3, a 2.8-Trillion-Parameter 'Open Frontier' Flagship — Weights Promised, Not Yet Downloadable
The strongest open-weight-class model yet shipped July 16 as a hosted service, with full weights committed for July 27 and no license published — buyers should treat 'open' as a dated promise, not an artifact.
Moonshot AI launched Kimi K3 on July 16, 2026, a 2.8-trillion-total-parameter Stable LatentMoE model with 1M-token context, native image and video input, and always-on reasoning, per launch coverage quoting Moonshot's announcement and technical blog.
Context
K3 is the closest an open-weight-class model has come to the frontier since DeepSeek R1, and the first open-announced model in the 3-trillion-parameter class. It landed in the middle of July's release wave, one day after Thinking Machines' Inkling and five days before Google's Gemini 3.6 Flash.
What changed
- Architecture: Stable LatentMoE with 16 of 896 experts active per token; Kimi Delta Attention with Attention Residuals; 1M-token context.
- Day-one availability: Kimi app and kimi.com (free tier), the `kimi-k3` OpenAI-compatible API, Kimi Work desktop, Kimi Code CLI, and OpenRouter.
- Pricing: $3 per million input tokens on cache miss, $0.30 on cache hit, $15 output.
- Variants: K3 Max and K3 Swarm Max.
- Weights: committed to public release by July 27, 2026 — but not downloadable at launch, and no license text, model card repo or checksums had been published as of July 22. K3 is API-only today.
Why it matters
K3 repriced the frontier-adjacent tier and demonstrated that a Chinese lab can ship near-frontier capability with open distribution intent. Moonshot's own technical blog concedes overall quality "still trails the strongest proprietary Claude Fable 5 and GPT-5.6 Sol models" while beating Opus 4.8 and GPT-5.5, per BreachRoad's enterprise analysis. Independent testing by Artificial Analysis placed K3 fourth of 189 models (score 57) and first on the Frontend Code Arena at 1,679 Elo — ahead of Fable 5 (1,631) and GPT-5.6 Sol (1,618) — with Vals.ai measuring 93.4% on SWE-bench Verified, per Vectrel's roundup.
Details
The launch also showcased agentic engineering feats — including a chip-design task completed in 48 hours at over 8,700 tokens/second throughput — as vendor demonstrations. Analyst Nathan Lambert called K3 "the closest open models have been to the frontier since DeepSeek R1." The gap between announcement and artifact matters: as aireiter's analysis notes, Moonshot's Hugging Face org still tops out at Kimi K2.7 Code.
Limitations and caveats
Do not cite a license. Press reports listing "MIT" or "Modified MIT" are inference from the K2-family precedent; no license text exists as of July 22. A Modified MIT license is widely expected based on K2/K2.7 precedent — whose attribution clause triggers at 100M monthly active users or $20M monthly revenue — but is unconfirmed. Launch-table benchmarks are vendor-reported; independent figures above come from third-party evaluators. Self-hosting will not relieve hosted demand soon: K3 requires 64+-accelerator supernodes.
Sources
- WccfTech — Kimi K3 launch coverage (quotes Moonshot announcement) (reputable press)
- BreachRoad — Kimi K3 architecture and enterprise security analysis (aggregator)
- Vectrel — Inkling, Kimi K3 and open-weight ownership strategy (aggregator)
- aireiter — Kimi K3 open-weights status (aggregator)
*Update note: This post was last reviewed on 2026-07-22. Next checkpoint: July 27 weights release — verify the repo on Moonshot's Hugging Face org, the license file text, and any MAU/revenue attribution clause.*
Sources
- WccfTech — Kimi K3 2.8T launch — reputable-press
- BreachRoad — Kimi K3 architecture/enterprise security — aggregator
- Vectrel — Inkling/Kimi K3 open-weight strategy — aggregator
- aireiter — Kimi K3 open weights — aggregator
Drafted with AI assistance from source briefs; reviewed for citation completeness and label accuracy.