On the BrowseComp benchmark - real-world research and browsing tasks - Kimi K3 scores 91.2, against Claude Fable 5 at 88.0 and GPT-5.6 Sol at 90.4. That is the first open-weight model to beat both closed flagships on that task, and it shipped with weights you can download. The catch is what "download" means here: Moonshot released Kimi K3's full weights on July 27, 2026, and the first thing most people notice is the download size - 1.56 TB on Hugging Face.
That single number is where most conversations about K3 should start, and mostly don't.
What Kimi K3 actually is
Kimi K3 is a 2.8T-parameter open-weight multimodal reasoning model from Moonshot AI, suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at navigating large repositories, using tools, debugging, and iterating against images, logs, tests, and runtime feedback.
The architecture is sparse. It uses a Mixture-of-Experts design with 896 experts, activating 16 per inference to maintain computational efficiency, and implements Kimi Delta Attention (KDA), a hybrid linear attention mechanism that provides a 6.3× decoding speedup in million-token contexts.
Kimi K3 is the first open model to reach 2.8 trillion parameters. For context: until K3, only closed models from OpenAI and Google DeepMind had publicly approached the trillion-parameter threshold. The MoE design keeps active compute manageable - roughly 104 billion parameters fire per token - but that does not shrink the weight file you need to load.
How K3 benchmarks against closed models
On the LLM-Stats composite as of early August 2026, Kimi K3 scores 55.4 against 57.2 for GPT-5.6 Sol and 56.5 for Claude Opus 5.
That gap is now smaller than the gap between the top two proprietary models.
Kimi K3 and GLM-5.3 tie on the Artificial Analysis Intelligence Index with a score of 60, just three points behind Claude Opus 5, and on DeepSWE, GLM-5.3 is the cheaper of the two to run at $3.99 per task against Kimi K3's $4.65.
The cost differential on coding is where K3's open pricing does real work. The rollout cost is $4.65 for Kimi K3 compared to $13.41 for Claude Fable 5. At a full 452-rollout sweep, the total cost is $2,103 versus $6,010.
That is a roughly 3× cost advantage on identical work. It does not require self-hosting to capture - OpenRouter lists $2.34 per million input tokens and $11.70 per million output tokens for the hosted API, with cached API input at $0.30 per million tokens as the headline rate.
One honest caveat: benchmark numbers need scrutiny. Benchmark questions leak into training corpora, and a model that has effectively memorized part of a test set will score well without generalizing. This is the main reason Humanity's Last Exam and SWE-bench Verified are weighted heavily by serious evaluators - HLE was built to be hard to saturate, and SWE-bench Verified was human-filtered specifically to remove tasks that were unsolvable or gameable.
The practical rule: use composites to decide which three models to try, and never to decide which one to ship.
Independent testing by Fireworks AI found specialization, not dominance: K3 won on security, crypto, and long terminal loops; Fable 5 won on multilingual and visualization tasks. Routing between them beat both.
| Dimension | Kimi K3 | Claude Fable 5 | GPT-5.6 Sol |
|---|---|---|---|
| LLM-Stats composite | 55.4 | 56.5 | 57.2 |
| BrowseComp | 91.2 | 88.0 | 90.4 |
| DeepSWE per-task cost | $4.65 | $13.41 | - |
| Weights available | Yes (July 27) | No | No |
| Context window | 1M tokens | - | - |
| Hosting | API + self-host | API only | API only |
The self-hosting math no one is running for you
This is the part most launch coverage skips. "Open-weight" does not mean runnable on ordinary hardware, and K3 is an extreme case.
Kimi K3 was trained quantization-aware and ships in MXFP4, which is already a 4-bit format. The usual trick of taking FP16 weights and quantizing them down to 4-bit to fit smaller hardware has already been spent. There is no easy 4x shrink left on the table, and community re-quantization below 4-bit typically costs real quality on a model like this. In other words, 1.56 TB is not the unoptimized size - it is the optimized size.
The checkpoint is 1,560.94 GB across 96 shards - about 160 GB more than 2.8T parameters at 4 bits predicts. That 160 GB difference is exactly the margin that decides whether an 8-GPU B200 node loads the model or dies at 94% of the way through weight loading, and it is the single most common sizing mistake on this model.
Moonshot recommends deploying Kimi K3 on large GPU clusters, with production deployments targeting supernode configurations of 64 or more accelerators.
The official vLLM recipe starts at 8 NVIDIA GB300 GPUs or 8 AMD MI355X/MI350X GPUs. That is one node minimum for a minimal setup, with a multi-node cluster with 2 TB or more of aggregate GPU memory needed for practical production use.
For most teams, the realistic K3 architecture in late 2026 is the routed one: API access for K3, open-weight deployment for smaller models like GLM-5.2 or Qwen 3.6 where the hardware math works, and a gateway in front of all of it.
Moonshot did contribute a vLLM implementation at release, so teams with the right infrastructure are not writing a custom serving stack. That is a genuine improvement over past large open-weight releases, but it does not change the GPU budget.
What to actually do with this information
Three decisions K3 clarifies right now:
- Test the API before touching the weights. At $2.34/$11.70 per million tokens on OpenRouter , you can run a representative sample of your real workload against K3 in an afternoon. That tells you whether the benchmark advantage shows up for your tasks before you think about infrastructure.
- Do not plan around further quantization. K3 shipped in MXFP4, which is already a 4-bit format. The usual post-release community quantization trick has already been spent. What you see is what you get.
- Route by task type, not model loyalty. The Fireworks finding - that routing between K3 and Fable 5 beat either model alone - is the most actionable result in K3's launch period. Neither open-weight nor closed wins on every axis, and the gap between them is small enough that mixing them by task type is worth the plumbing.
Kimi K3 open-weight model: common questions
What is Kimi K3 and who built it?
Kimi K3 is Moonshot AI's flagship multimodal reasoning model, released as a hosted service on July 16, 2026. Its 2.8 trillion total parameters make it the first announced open model in the three-trillion-parameter class. Moonshot is a Chinese AI lab, and K3 ships under a custom Kimi K3 License for commercial use.
How does Kimi K3 benchmark against GPT-5.6 and Claude?
On the LLM-Stats composite as of early August 2026, Kimi K3 scores 55.4 against 57.2 for GPT-5.6 Sol and 56.5 for Claude Opus 5. It ranks fifth overall. That is a difference of 1.8 points between the best model you can download and the best model you can only rent. On BrowseComp it leads both.
Can you self-host Kimi K3?
Yes, but the hardware bar is high. The official vLLM recipe starts at 8 NVIDIA GB300 GPUs or 8 AMD MI355X/MI350X GPUs , and Moonshot's production guidance targets 64 or more accelerators. The weights are 1.56 TB in already-optimized MXFP4 format, so there is no simple way to shrink them further for smaller hardware.
What does Kimi K3 cost on the API?
OpenRouter lists $2.34 per million input tokens and $11.70 per million output tokens, with a 1,048,576-token context window.
Cached input drops to $0.30 per million tokens. That is roughly a third of the per-task cost on Claude Fable 5 for comparable coding workloads.
Is Kimi K3's MXFP4 quantization different from post-release community quants?
Yes, meaningfully so. MXFP4 is not a post-hoc quantization applied after training. K3 was trained with quantization-aware training from the SFT stage onward, so the released MXFP4 weights serve natively at 4-bit without a separate calibration pass. That is a better starting point than most community quants, and it also means you cannot shrink the model further without real quality loss.