Moonshot AI dropped 96 Hugging Face shards on July 27, 2026 - 1.56 terabytes of MXFP4 weights that together make up Kimi K3, the largest open-weight model ever released. The headline number is 2.8 trillion parameters. The practical question is whether that release means anything for your team, or whether it is open in the same way a concert is open when you can watch it through a window from across the street.
The answer depends on which of three very different use cases you are in: API consumer, self-hosting operator, or enterprise buyer thinking about compliance.
What Kimi K3 actually is
Kimi K3 is a 2.8-trillion-parameter open-weight multimodal reasoning model from Moonshot AI. That number is attention-grabbing, but the architecture is why it is not as absurd as it sounds. It uses a Mixture-of-Experts architecture with 896 experts, activating just 16 per inference to maintain computational efficiency.
That 1.8% activation ratio means each token touches roughly 104 billion parameters out of 2.8 trillion. So compute per token is comparable to a dense ~100B model - it is the memory footprint that is enormous.
The key architectural innovation for long-context work is Kimi Delta Attention (KDA), a hybrid linear attention mechanism that provides a 6.3× decoding speedup in million-token contexts.
It also requires a modified prefix caching implementation, and Moonshot contributed a KDA-compatible prefill caching build to vLLM alongside the weights release. That is genuinely useful: it means the 1M-token context window is not just a spec-sheet number but is served with a proper caching path.
Moonshot AI publicly released Kimi K3 on July 16, 2026, with full open-source weights released by July 27. At 2.8 trillion parameters, it is the first open-source model to reach the 3-trillion-parameter class.
On benchmarks, on BrowseComp - which evaluates real-world research and browsing - Kimi K3 scores 91.2, compared to Claude Fable 5 at 88.0 and GPT-5.6 Sol at 90.4.
On DeepSWE with the mini-SWE-agent harness, K3 achieves 67.3 percent.
The harness caveat is real: different evaluation harnesses can swing coding scores by 10 to 26 points. A model that leads one leaderboard is not automatically leading your actual workflow.
What it actually costs to use Kimi K3
Kimi K3 costs $3.00 per million cache-miss input tokens and $15.00 per million output tokens, with a $0.30 cache-hit input rate and a 1M-token context window. Compare that to its predecessor: K3 is priced roughly 3.5-4× above K2.6/K2.7 Code's rate, on a context window 4× larger than K2's 256K.
The cache discount is the most underreported part of this pricing. With a 1M-token context window, long shared prefixes are common in agentic and RAG workflows
- meaning any team running a long system prompt repeatedly will see a large fraction of their input tokens hit the cache at $0.30, not $3.00. On a coding agent that embeds the same 200K-token codebase context in every call, the effective input price could be closer to $0.30 than $3.00. Do the math for your specific call pattern before assuming the headline rate applies.
Self-hosting is a different conversation. The weights alone need roughly 1.4 TB of VRAM at the native 4-bit quantisation, which in practice means an 8× NVIDIA B200 node (1,536 GB) at the low end, or 16× H200 across two nodes for the full 1-million-token context.
At a mid-market rate near $5.50 per GPU-hour, an 8× B200 node runs about $32,000 a month - so the break-even against the $15/M API sits above 2 billion output tokens a month. Below that volume, the managed API wins on both price and operational effort.
| Route | Input ($/1M) | Output ($/1M) | Min monthly infra | Break-even volume |
|---|---|---|---|---|
| Moonshot API (cache-miss) | $3.00 | $15.00 | $0 | - |
| Moonshot API (cache-hit) | $0.30 | $15.00 | $0 | - |
| Self-host (8× B200, cloud) | $0 token fee | $0 token fee | ~$32,000 | ~2.1B output tokens/mo |
| Self-host (8× B300, cloud) | $0 token fee | $0 token fee | ~$45,000 | - |
The cheapest rental that fits is 8× B300 at about $45,331 a month - the same as roughly 9 billion tokens a month on Moonshot AI's API. Below that volume, self-hosting is more expensive. For most engineering teams, the answer is: use the API.
The compliance gap nobody is talking about
This is the part that most Kimi K3 coverage skips entirely.
Moonshot stores API data in Singapore, trains on your content by default with no documented opt-out, and doesn't publish a SOC 2 report or offer a signed HIPAA BAA. For a US engineering team running code review or debugging through the API, that is a policy conversation, not a technical one. For a healthcare or financial services team, it may be a hard blocker.
The license has teeth at scale too. Two conditions apply only at scale: model-as-a-service businesses earning over $20M a year on K3 need a separate agreement with Moonshot, and products with over 100M monthly users or $20M in monthly revenue must display "Kimi K3" in their interface. Neither clause matters for most enterprise internal tooling, but they matter a lot if you are building a product on top of K3.
The self-hosting path bypasses the data residency issue - your tokens never leave your cluster - but brings the infrastructure burden described above. A teammate like Beagle, running inside Slack or Teams, would still need to route requests through either Moonshot's API or a self-hosted endpoint you control, so the compliance question lands at your infra layer, not the application layer.
What is genuinely new versus incremental
Kimi K3 introduces several architectural innovations that collectively yield approximately 2.5× scaling efficiency improvement over its predecessor K2. The KDA mechanism and quantization-aware training from the SFT stage onward - using MXFP4 weights with MXFP8 activations for broad hardware compatibility
- are real advances, not marketing.
What is incremental: the benchmark lead. The composite picture at the top of the leaderboard is that Claude Fable 5.1 sits first, but the entire top five spans 1.6 points while covering a 119× price range. K3 is inside that 1.6-point band. Independent testing by Fireworks AI found specialization, not dominance: K3 won on security, crypto, and long terminal loops; Fable 5 won on multilingual and visualization tasks. Routing between them beat both.
What is genuinely new: the open-weight milestone itself. Kimi K3 and GLM-5.3 tie on the Artificial Analysis Intelligence Index with a score of 60, just three points behind Claude Opus
- and K3 ships real weights you can download and inspect. That matters for auditing, fine-tuning, and research in a way that a closed API does not, regardless of benchmark rank.
Kimi K3 open weight model: common questions
What hardware does Kimi K3 require to self-host?
The minimum production configuration is an 8× NVIDIA B200 node with 1,536 GB of HBM3e, or AMD MI355X/MI350X equivalents. A single 8× H100 node - at 640 GB - falls short of the 1.4 TB VRAM floor. For the full 1M-token context window, a 16× H200 two-node setup is the recommended path.
How does Kimi K3 pricing compare to its predecessors?
Kimi K3 costs $3.00 per million cache-miss input tokens and $15.00 per million output tokens, with a $0.30 cache-hit input rate. The K2.6 and K2.7 Code tier costs $0.95 input / $4.00 output. That is roughly a 3.75× step up in price, paired with a 4× larger context window and the KDA architecture for long-context throughput.
Is Kimi K3 actually open source?
The weights are real and downloadable. The weights were released July 27, 2026 under the Kimi K3 License. It is free to use, modify, and deploy commercially with attribution. It is not Apache 2.0: there are revenue and user-count thresholds above which additional terms apply. "Open weight" is accurate; "fully open source" is a stretch.
How does Kimi K3 perform on coding benchmarks?
On DeepSWE with the mini-SWE-agent harness, K3 achieves 67.3 percent. The harness caveat is real: different evaluation harnesses can swing coding scores by 10 to 26 points. Treat published numbers as a filter for a short list, then run your own eval on a sample of your actual codebase before committing.
Can enterprise teams use the Kimi K3 API?
Not without a compliance conversation first. Moonshot stores API data in Singapore, has no documented opt-out from training on API content, and publishes no SOC 2 report or HIPAA BAA. Teams with data residency requirements or regulated data should either negotiate a DPA with Moonshot directly or deploy self-hosted weights in their own environment.