Kimi K3: 2.8 Trillion Parameters, Open Weights, and a New AI Frontier
- Kimi K3: The 2.8 Trillion-Parameter Open-Weight Giant
- Core Specs: What Makes K3 Different
- Architecture Deep Dive: Delta Attention & MoE
- Benchmarks: How K3 Stacks Up
- Head-to-Head: K3 vs GPT-5.6 vs Claude Fable 5 vs GLM 5.2
- Pricing & Access: What It Costs
- Open Weights: The July 27 Milestone
- Why K3 Matters for the AI Industry
- Key Takeaways
July 22, 2026 — On July 16, Moonshot AI dropped what may be the most significant AI model launch of the year. Kimi K3 is not just another incremental upgrade. At 2.8 trillion parameters, it is the largest open-weight model ever announced, and independent benchmarks place it fourth globally — within striking distance of OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 [citation:1][citation:4].
What makes K3 remarkable is not just its size. It is the first Chinese model to credibly challenge the American frontier on capability rather than price alone. And it is doing so with its weights set to go fully open on July 27 — a move that could reshape how enterprises and developers choose their AI infrastructure [citation:3][citation:5].
Core Specs: What Makes K3 Different
K3's headline numbers are striking, but the architecture behind them is what matters:
| Specification | Kimi K3 |
|---|---|
| Total parameters | 2.8 trillion |
| Architecture | Mixture-of-Experts (896 experts, 16 active per token) |
| Context window | 1,000,000 tokens |
| Attention mechanism | Kimi Delta Attention (KDA) + Attention Residuals |
| Multimodal | Native image and video understanding |
| Reasoning | Always-on "thinking mode" |
| Decoding speed | ~6.3x faster long-context decoding vs prior gen |
| Output speed | 62 tokens/second (independent measurement) |
| Open weights | Scheduled July 27, 2026, Modified MIT license |
| API pricing | $3.00 in / $15.00 out per million tokens |
The 2.8 trillion parameter count is the headline, but the practical number is the active compute. With only 16 of 896 experts activated per token, K3 delivers frontier-level performance without requiring frontier-level inference costs [citation:2][citation:5].
Architecture Deep Dive: Delta Attention & MoE
Moonshot developed two in-house architectural innovations for K3:
Kimi Delta Attention (KDA)
A hybrid linear-attention mechanism designed to solve the long-context bottleneck. Standard transformers slow dramatically as context length grows. KDA reportedly delivers 6.3x faster decoding at 1 million tokens, making the full context window usable in practice rather than just a spec-sheet number [citation:5][citation:6].
Attention Residuals
Described as a drop-in replacement for standard residual connections. Moonshot claims a 25% training-efficiency gain for under 2% additional compute overhead — a critical advantage when training a 2.8T model under China's compute constraints [citation:5].
Stable LatentMoE
K3's mixture-of-experts framework activates just 16 of 896 experts per forward pass. This sparse design keeps inference costs manageable despite the massive parameter count, converting what would be an impractical model into a deployable one [citation:2][citation:6].
Benchmarks: How K3 Stacks Up
Independent testing gives a clearer picture than launch-day marketing:
| Benchmark | Kimi K3 | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|---|
| Artificial Analysis Intelligence Index | 57 (#4) | ~59 (#2) | ~60 (#1) |
| Terminal-Bench 2.1 (coding) | 88.3 | 88.8 | 88.0 |
| Frontend Code Arena | #1 | #2 | #3 |
| 1M-token context (no management) | 90.4 | N/A | N/A |
| Output tokens/second | 62 | Varies | Varies |
The takeaway: K3 does not top every leaderboard, but it is competitive everywhere that matters. It leads on frontend coding, holds its own on general reasoning, and delivers the best long-context performance of any open-weight model [citation:4][citation:6][citation:7].
Head-to-Head: K3 vs GPT-5.6 vs Claude Fable 5 vs GLM 5.2
| Dimension | Kimi K3 | GPT-5.6 Sol | Claude Fable 5 |
|---|---|---|---|
| Parameters | 2.8T (16 active) | Undisclosed | Undisclosed |
| Context window | 1M tokens | Up to 272K | Up to 200K |
| Open weights | Yes (July 27) | No | No |
| Input price | $3.00/M | $5.00/M | Higher tier |
| Output price | $15.00/M | $30.00/M | Higher tier |
| Multimodal | Native image + video | Yes | Yes |
| Coding rank | #1 (Frontend) | #2 | #3 |
| Overall rank | #4 | #2 | #1 |
| License | Modified MIT | Proprietary | Proprietary |
Against GLM 5.2 — Zhipu AI's coding-focused flagship — the comparison is different:
| Dimension | Kimi K3 | GLM 5.2 |
|---|---|---|
| Parameters | 2.8T | ~744B |
| Design goal | Frontier general intelligence | Coding/agent optimization |
| Context window | 1M tokens | 1M tokens |
| Input price | $3.00/M | $1.40/M |
| Output price | $15.00/M | $4.40/M |
| Intelligence Index | 57 | 51 |
| Best for | General reasoning + coding | Pure coding workloads |
Pricing & Access: What It Costs
K3 launched with full availability across Moonshot's product suite:
- Kimi app & kimi.com — Free tier for everyone (iOS, Android, HarmonyOS, web)
- Kimi Work — Desktop app for Windows and Apple Silicon
- Kimi API — OpenAI-compatible endpoints at platform.moonshot.ai
- Kimi Code — Terminal/IDE agent with K3 selectable via /model
| Tier | Price | Notes |
|---|---|---|
| Input (cache miss) | $3.00 / million tokens | Standard API rate |
| Input (cache hit) | $0.30 / million tokens | Repeated context |
| Output | $15.00 / million tokens | No premium for 1M context |
| Free tier | $0 | Via Kimi app and kimi.com |
Moonshot also launched a referral program: new users who sign up through an invite link receive bonus membership credits — up to a full year free [citation:1].
Open Weights: The July 27 Milestone
The most consequential date on the calendar is July 27, 2026. That is when Moonshot has committed to releasing K3's full weights under a Modified MIT license [citation:1][citation:5].
If that happens, K3 will become:
- The largest open-weight model ever released — nearly 3x the size of K2's 1.0T
- The first open model in the 3T-parameter class
- A viable self-hosting option for enterprises with sufficient GPU infrastructure
Why K3 Matters for the AI Industry
K3's launch has already sent ripples through the market:
- Chip stocks wobbled — Investors fear open Chinese models could accelerate AI commoditization and reduce demand for high-end training chips [citation:6]
- Enterprise negotiating power shifts — A frontier-capable open model gives companies leverage against OpenAI and Anthropic pricing [citation:4]
- China's AI credibility rises — K3 is the first Chinese model to compete on capability, not just cost [citation:3][citation:6]
- Open vs closed debate intensifies — If a 2.8T open model can match closed frontiers, the case for proprietary lock-in weakens [citation:5]
Moonshot is also reportedly preparing for a Hong Kong IPO, with a valuation of approximately $31.5 billion. If successful, it would be the first Chinese large-model unicorn to go public [citation:6].
Key Takeaways
| # | What You Need to Know About Kimi K3 |
|---|---|
| 1 | 2.8 trillion parameters — the largest open-weight model ever announced, using Stable LatentMoE with 896 experts and 16 active per token [citation:1][citation:5] |
| 2 | Ranks #4 globally on the Artificial Analysis Intelligence Index, trailing only Claude Fable 5, GPT-5.6 Sol, and GPT-5.6 Sol Ultra [citation:4][citation:7] |
| 3 | Leads Frontend Code Arena — ahead of both Claude Fable 5 and GPT-5.6 Sol on coding benchmarks [citation:6][citation:7] |
| 4 | 1 million token context window with 6.3x faster long-context decoding via Kimi Delta Attention [citation:5] |
| 5 | API pricing is competitive — $3 in / $15 out per million tokens, roughly half the output cost of GPT-5.6 [citation:7] |
| 6 | Open weights drop July 27, 2026 under Modified MIT license — the largest open release in history [citation:1][citation:5] |
| 7 | Available now via Kimi app, Kimi Work desktop, Kimi Code CLI, and OpenAI-compatible API [citation:1] |
| 8 | Native multimodal — image and video understanding built in, not bolted on [citation:5] |
- K3-kimi.com — Official release timeline and access channels [citation:1]
- Falconer — Kimi K2 vs K3 architecture comparison [citation:2]
- Collective Brain — Kimi K3 business guide and enterprise analysis [citation:3]
- Codelabra — K3 benchmarks, Terminal-Bench scores, and industry impact [citation:4][citation:6]
- Codersera — K3 complete guide, specs, and pricing [citation:5]
- Dealcan / Forbes — Market impact, IPO preparation, and chip stock reaction [citation:6]
- Gamine AI — K3 review, API value analysis, and GPT-5.6 comparison [citation:7]
- Planet Tools AI — GLM 5.2 vs K3 pricing and capability comparison [citation:8]
From high-performance GPUs for model training to workstations optimized for AI development — we have the hardware to match your ambition.
Special Offer: Use code KIMIK3 for a discount on AI hardware and accessories!
Shop AI Hardware at GzmatoNvidia RTX, AMD Radeon, server GPUs, and more — shipped worldwide.
Our team can recommend the right GPUs, workstations, and cloud setups for running or fine-tuning models like Kimi K3. Chat with us or open a request.
- Kimi K3
- Moonshot AI
- Kimi K3 review
- Kimi K3 benchmarks
- Kimi K3 vs GPT-5.6
- Kimi K3 vs Claude
- open weight AI model
- 2.8 trillion parameters
- Kimi K3 pricing
- Kimi K3 API
- Chinese AI model 2026
- GLM 5.2 vs Kimi K3
