July 22, 2026 — On July 16, Moonshot AI dropped what may be the most significant AI model launch of the year. Kimi K3 is not just another incremental upgrade. At 2.8 trillion parameters, it is the largest open-weight model ever announced, and independent benchmarks place it fourth globally — within striking distance of OpenAI's GPT-5.6 Sol and Anthropic's Claude Fable 5 [citation:1][citation:4].

What makes K3 remarkable is not just its size. It is the first Chinese model to credibly challenge the American frontier on capability rather than price alone. And it is doing so with its weights set to go fully open on July 27 — a move that could reshape how enterprises and developers choose their AI infrastructure [citation:3][citation:5].

Quick Answer: Kimi K3 is Moonshot AI's 2.8-trillion-parameter flagship, released July 16, 2026. It ranks #4 globally on independent benchmarks, leads the Frontend Code Arena, and carries a 1-million-token context window. Open weights drop July 27 under a Modified MIT license. For developers and enterprises, K3 represents the first open-weight model that does not require a capability compromise versus closed American frontiers.

Core Specs: What Makes K3 Different

K3's headline numbers are striking, but the architecture behind them is what matters:

Specification Kimi K3
Total parameters2.8 trillion
ArchitectureMixture-of-Experts (896 experts, 16 active per token)
Context window1,000,000 tokens
Attention mechanismKimi Delta Attention (KDA) + Attention Residuals
MultimodalNative image and video understanding
ReasoningAlways-on "thinking mode"
Decoding speed~6.3x faster long-context decoding vs prior gen
Output speed62 tokens/second (independent measurement)
Open weightsScheduled July 27, 2026, Modified MIT license
API pricing$3.00 in / $15.00 out per million tokens

The 2.8 trillion parameter count is the headline, but the practical number is the active compute. With only 16 of 896 experts activated per token, K3 delivers frontier-level performance without requiring frontier-level inference costs [citation:2][citation:5].


Architecture Deep Dive: Delta Attention & MoE

Moonshot developed two in-house architectural innovations for K3:

Kimi Delta Attention (KDA)

A hybrid linear-attention mechanism designed to solve the long-context bottleneck. Standard transformers slow dramatically as context length grows. KDA reportedly delivers 6.3x faster decoding at 1 million tokens, making the full context window usable in practice rather than just a spec-sheet number [citation:5][citation:6].

Attention Residuals

Described as a drop-in replacement for standard residual connections. Moonshot claims a 25% training-efficiency gain for under 2% additional compute overhead — a critical advantage when training a 2.8T model under China's compute constraints [citation:5].

Stable LatentMoE

K3's mixture-of-experts framework activates just 16 of 896 experts per forward pass. This sparse design keeps inference costs manageable despite the massive parameter count, converting what would be an impractical model into a deployable one [citation:2][citation:6].

Why this matters: The combination of KDA + Attention Residuals + Stable LatentMoE is what lets a 2.8T model run at prices comparable to mid-range Western models. The architecture, not just scale, is the story.

Benchmarks: How K3 Stacks Up

Independent testing gives a clearer picture than launch-day marketing:

Benchmark Kimi K3 GPT-5.6 Sol Claude Fable 5
Artificial Analysis Intelligence Index57 (#4)~59 (#2)~60 (#1)
Terminal-Bench 2.1 (coding)88.388.888.0
Frontend Code Arena#1#2#3
1M-token context (no management)90.4N/AN/A
Output tokens/second62VariesVaries

The takeaway: K3 does not top every leaderboard, but it is competitive everywhere that matters. It leads on frontend coding, holds its own on general reasoning, and delivers the best long-context performance of any open-weight model [citation:4][citation:6][citation:7].


Head-to-Head: K3 vs GPT-5.6 vs Claude Fable 5 vs GLM 5.2

Dimension Kimi K3 GPT-5.6 Sol Claude Fable 5
Parameters2.8T (16 active)UndisclosedUndisclosed
Context window1M tokensUp to 272KUp to 200K
Open weightsYes (July 27)NoNo
Input price$3.00/M$5.00/MHigher tier
Output price$15.00/M$30.00/MHigher tier
MultimodalNative image + videoYesYes
Coding rank#1 (Frontend)#2#3
Overall rank#4#2#1
LicenseModified MITProprietaryProprietary

Against GLM 5.2 — Zhipu AI's coding-focused flagship — the comparison is different:

Dimension Kimi K3 GLM 5.2
Parameters2.8T~744B
Design goalFrontier general intelligenceCoding/agent optimization
Context window1M tokens1M tokens
Input price$3.00/M$1.40/M
Output price$15.00/M$4.40/M
Intelligence Index5751
Best forGeneral reasoning + codingPure coding workloads
The price advantage is real but nuanced: K3's input price is 40% cheaper than GPT-5.6, and its output price is 50% cheaper. But K3 tends to use more output tokens per task, which can narrow the gap. For cost-sensitive workloads, GLM 5.2 remains the budget king [citation:7][citation:8].

Pricing & Access: What It Costs

K3 launched with full availability across Moonshot's product suite:

  • Kimi app & kimi.com — Free tier for everyone (iOS, Android, HarmonyOS, web)
  • Kimi Work — Desktop app for Windows and Apple Silicon
  • Kimi API — OpenAI-compatible endpoints at platform.moonshot.ai
  • Kimi Code — Terminal/IDE agent with K3 selectable via /model
Tier Price Notes
Input (cache miss)$3.00 / million tokensStandard API rate
Input (cache hit)$0.30 / million tokensRepeated context
Output$15.00 / million tokensNo premium for 1M context
Free tier$0Via Kimi app and kimi.com

Moonshot also launched a referral program: new users who sign up through an invite link receive bonus membership credits — up to a full year free [citation:1].


Open Weights: The July 27 Milestone

The most consequential date on the calendar is July 27, 2026. That is when Moonshot has committed to releasing K3's full weights under a Modified MIT license [citation:1][citation:5].

If that happens, K3 will become:

  • The largest open-weight model ever released — nearly 3x the size of K2's 1.0T
  • The first open model in the 3T-parameter class
  • A viable self-hosting option for enterprises with sufficient GPU infrastructure
Reality check: A 2.8T MoE model is not trivial to self-host. Even with sparse activation, the hardware requirements are substantial. Most teams will use the hosted API. The open weights matter most for researchers, fine-tuners, and enterprises with existing GPU clusters [citation:5][citation:7].

Why K3 Matters for the AI Industry

K3's launch has already sent ripples through the market:

  • Chip stocks wobbled — Investors fear open Chinese models could accelerate AI commoditization and reduce demand for high-end training chips [citation:6]
  • Enterprise negotiating power shifts — A frontier-capable open model gives companies leverage against OpenAI and Anthropic pricing [citation:4]
  • China's AI credibility rises — K3 is the first Chinese model to compete on capability, not just cost [citation:3][citation:6]
  • Open vs closed debate intensifies — If a 2.8T open model can match closed frontiers, the case for proprietary lock-in weakens [citation:5]

Moonshot is also reportedly preparing for a Hong Kong IPO, with a valuation of approximately $31.5 billion. If successful, it would be the first Chinese large-model unicorn to go public [citation:6].


Key Takeaways

# What You Need to Know About Kimi K3
12.8 trillion parameters — the largest open-weight model ever announced, using Stable LatentMoE with 896 experts and 16 active per token [citation:1][citation:5]
2Ranks #4 globally on the Artificial Analysis Intelligence Index, trailing only Claude Fable 5, GPT-5.6 Sol, and GPT-5.6 Sol Ultra [citation:4][citation:7]
3Leads Frontend Code Arena — ahead of both Claude Fable 5 and GPT-5.6 Sol on coding benchmarks [citation:6][citation:7]
41 million token context window with 6.3x faster long-context decoding via Kimi Delta Attention [citation:5]
5API pricing is competitive — $3 in / $15 out per million tokens, roughly half the output cost of GPT-5.6 [citation:7]
6Open weights drop July 27, 2026 under Modified MIT license — the largest open release in history [citation:1][citation:5]
7Available now via Kimi app, Kimi Work desktop, Kimi Code CLI, and OpenAI-compatible API [citation:1]
8Native multimodal — image and video understanding built in, not bolted on [citation:5]
Kimi K3 is the moment open weights caught up to the closed frontier. It will not win every benchmark, and self-hosting a 2.8T model is not practical for most teams. But for the first time, developers and enterprises can reach for an open model without accepting a capability gap. That changes the math for every AI procurement decision going forward.
Sources and Methodology (as of July 22, 2026):
  • K3-kimi.com — Official release timeline and access channels [citation:1]
  • Falconer — Kimi K2 vs K3 architecture comparison [citation:2]
  • Collective Brain — Kimi K3 business guide and enterprise analysis [citation:3]
  • Codelabra — K3 benchmarks, Terminal-Bench scores, and industry impact [citation:4][citation:6]
  • Codersera — K3 complete guide, specs, and pricing [citation:5]
  • Dealcan / Forbes — Market impact, IPO preparation, and chip stock reaction [citation:6]
  • Gamine AI — K3 review, API value analysis, and GPT-5.6 comparison [citation:7]
  • Planet Tools AI — GLM 5.2 vs K3 pricing and capability comparison [citation:8]
Published: July 22, 2026. Kimi K3 launched July 16, 2026. Open weights are scheduled for release July 27, 2026, under a Modified MIT license. Benchmark rankings are based on independent third-party testing and may shift as more evaluations are completed.
Power Your AI Workflow with Gzmato

From high-performance GPUs for model training to workstations optimized for AI development — we have the hardware to match your ambition.

Special Offer: Use code KIMIK3 for a discount on AI hardware and accessories!

Shop AI Hardware at Gzmato

Nvidia RTX, AMD Radeon, server GPUs, and more — shipped worldwide.

Need help choosing AI hardware for your workload?

Our team can recommend the right GPUs, workstations, and cloud setups for running or fine-tuning models like Kimi K3. Chat with us or open a request.