On Monday 27 July, Moonshot AI published the weights of Kimi K3 on Hugging Face, ten days after announcing the model on 17 July and on the day it had promised. The checkpoint is about 1.56 terabytes. The model is 2.8 trillion parameters, 104 billion of them active for any one token, with a one-million-token context and native vision. It ships under a five-clause document called the Kimi K3 License, which permits commercial use, asks nothing of anyone who runs it internally, and asks large hosting businesses for a separate agreement. By the evening, OpenRouter listed it from seven providers at the $3 per million input tokens and $15 per million output tokens that Moonshot charges itself. Moonshot’s price list carries a third number that matters more than either: $0.30 per million for input served from cache.
By its publisher’s benchmarks this is the first open-weight model at the frontier, and it arrived with a price list built around the cache rather than the token. For an enterprise the question is not whether the model is good. It is whether to hold the weights, and that depends on arithmetic almost no team has done: what it costs to keep 2.8 trillion parameters warm, what utilisation you need before that beats a $0.30 cache hit, and what you inherit when there is no vendor between you and the model. The arithmetic is below, every assumption labelled, and it is less flattering to self-hosting than the word “open” suggests.
What shipped, and the number that matters
In plain terms, K3 is a mixture-of-experts transformer. Each layer holds 896 expert networks, a router picks 16 of them for every token, and two shared experts see every token. That is how a 2.8-trillion-parameter model does the work of one a twenty-seventh of its size on each token: 104 billion parameters are active and the rest wait in memory. The model is 93 layers deep. Sixty-nine use Kimi Delta Attention, Moonshot’s linear-attention design, which keeps a fixed-size state rather than a key-value cache that grows with every token; 24 use Gated MLA, a gated form of the latent attention that compresses the key-value cache. That ratio is what makes a million-token context affordable to serve. Attention Residuals, the third name on the list, changes how information passes between layers; the technical report has the detail. Vision is native, through a 401-million-parameter encoder for images and video.
The precision is the part to read twice. Weights are MXFP4, a four-bit floating-point format with per-block scaling; activations are MXFP8; and Moonshot did quantisation-aware training from the supervised fine-tuning stage onward, so the four-bit weights are the model, not a lossy export of it. Two things follow. First, 2.8 trillion parameters at half a byte each is about 1.4 terabytes of weight data; with embeddings, the vision encoder and scaling factors, the checkpoint is about 1.56 terabytes. Second, MXFP4 runs natively on NVIDIA Blackwell and AMD MI400; on the Hopper generation most enterprises own, the weights would be unpacked into a wider format, roughly doubling the memory.
Why the active count matters for serving: compute per token scales with the 104 billion active parameters, memory with the 2.8 trillion total. Per token, K3 costs about what a 100-billion-parameter dense model costs to run, which is why Moonshot can sell output at $15 per million. Per deployment, it costs what a 2.8-trillion-parameter model costs to hold, which is why nobody runs it on one machine. The total sets the floor on hardware; the active count sets what you get once you have paid for it.
Moonshot reports 88.3 on Terminal-Bench 2.1, 67.5 on DeepSWE, 93.5 on GPQA Diamond and 91.2 on BrowseComp at maximum reasoning effort. Those are the publisher’s numbers, and the bench in the evaluation post exists because publisher numbers begin a qualification rather than end it.
The licence, clause by clause
The Kimi K3 License is 488 words long, a custom document rather than Apache or MIT, and custom documents are where open-weight projects stall. Here it is clause by clause, to be read with counsel.
The grant. MIT in form: permission, free of charge, to use, copy, modify, merge, publish, distribute, sublicense and sell the software, which includes weights, code and documentation, and to run, deploy, fine-tune and create derivative works. No payment appears anywhere.
Clause 1. The notice travels with all copies, and use must comply with applicable law. For most users that is the entire obligation.
Clause 2. The commercial gate, with two conditions that must both hold. Model as a Service means giving a third party access to inference or fine-tuning with meaningful control over inputs, parameters or training data; products with the model embedded in a feature, and the relaying of requests to a model hosted elsewhere, are excluded. If the licensee or any of its affiliates operates such a business, and the aggregate revenue of the licensee and its affiliates exceeds $20 million over any consecutive twelve months, the licensee must enter a separate agreement with Moonshot before using the software or its derivatives for any commercial purpose. Note what the revenue test measures: not the hosting business, but the group. A conglomerate with a cloud subsidiary selling hosted inference anywhere in its structure meets both conditions at any revenue a listed company has. The consequence is not a fee but a negotiation that must precede any commercial use.
Clause 3. A commercial product with more than 100 million monthly active users, or more than $20 million in monthly revenue, must display “Kimi K3” prominently on its interface. No fee, but for a bank or a telecom whose app clears either threshold, a model name on the screen is a brand decision and possibly a regulatory one.
Clause 4. The exemptions. Clauses 2 and 3 do not apply to internal use, defined as use that “does not make the Software, its outputs, or its underlying capabilities available to third parties”, nor to use through Moonshot’s own products or certified inference partners. That definition decides most enterprise cases. A coding assistant for your engineers, a document pipeline, a back-office classifier, an assistant for your own contact-centre staff: none makes the model or its outputs available to third parties, so they owe nothing and display nothing. A customer-facing chatbot does, so it is not internal use; it then meets clause 3 only above the thresholds and clause 2 only if the group runs a hosting business.
Clause 5. The software, and expressly its outputs, come as is, without warranty or liability. A hosted provider’s enterprise contract usually carries an SLA and some indemnity; here there is neither.
What the document omits is as informative. No governing law, so a dispute falls to whatever conflict-of-laws rules apply, with a licensor in Beijing. No patent grant, which Apache 2.0 has; no acceptable-use policy, no field-of-use restriction, no termination clause, no obligation to share improvements and no telemetry. Simon Willison noted on the day that Moonshot says “open weight” and never “open source”, and that is the right reading; the July comparison is Thinking Machines’ Inkling, 975 billion parameters released on 15 July under Apache 2.0, which a legal team clears without reading. For an Indian enterprise K3’s terms come down to three questions: is the use internal, does any affiliate sell hosted inference, and will the product cross 100 million users or $20 million a month? The answers belong in the model registry beside the router configuration described in the routing post, so that no workload can be pointed at the model under terms that exclude it.
What it costs to hold 2.8 trillion parameters
Memory first. An HGX B200 node carries eight GPUs with 180 gigabytes each, 1.4 terabytes in total; the checkpoint is 1.56 terabytes, so one node does not hold the model. Two nodes, sixteen GPUs and 2.88 terabytes, hold the weights with about 1.3 terabytes left for key-value cache and activations. That is the floor that fits, not the floor that performs: each GPU holds 56 experts of every layer, the router’s all-to-all traffic crosses the inter-node link on every token, and a million-token context for a few dozen concurrent users eats the spare memory quickly. Moonshot recommends a supernode of 64 or more accelerators in one interconnect domain: a GB200 NVL72 rack, 72 GPUs and 13.4 terabytes, or eight HGX nodes with expert traffic over InfiniBand. The smallest community quantisation, Unsloth’s one-bit dynamic GGUF, is 594 gigabytes, needs about 610 gigabytes of combined memory, agrees with the full model on 78.9 percent of top-1 predictions, and runs at roughly 20 tokens per second on a B200: usable for a lab, not for two hundred engineers.
Then the rent. By a July survey of rental rates, B200s run between $4.99 and $6.04 per GPU-hour on demand at specialist clouds, 36-month reserved contracts go as low as $2.25, and a GB200 NVL72 rack prices between $756 and $1,944 an hour. At $5.50 and 730 hours a month, the sixteen-GPU floor is about $64,000 a month and the 64-GPU floor about $257,000; reserved for three years at $2.25, the 64-GPU cluster is about $105,000. The rack is $552,000 to $1.42 million a month.
Then the people. Serving this model means expert-parallel inference across nodes, kernels for Kimi Delta Attention, an MXFP4 path that preserves quality, a prefix cache tier, observability, patching the stack every few weeks and on-call for a tier-one dependency. Moonshot publishes recipes for vLLM, SGLang and its own TokenSpeed; someone on your payroll still has to own them. I budget $30,000 a month for a three-person serving team with on-call cover: my estimate at Bengaluru rates, and the cheapest line in the sum.
Now the break-even. Self-hosting is a fixed cost, hardware plus team; the API is a variable cost, tokens times a blended price. The crossover is fixed cost divided by blended price, in tokens, and the utilisation you need is that count divided by the cluster’s capacity. The fixed cost is above. The blended price comes from the worked example below: at a coding agent’s token mix, with 90 percent of input from cache, Kimi’s API costs about $0.85 per million workload tokens and Claude Opus 5’s about $1.54. The capacity is the number nobody has published. I assume a 64-GPU cluster serves 400 billion workload tokens a month at full load in that mix, conservative since nine tokens in ten are cache hits that never touch a tensor core, and an estimate. On those assumptions the reserved cluster at $135,000 a month beats Opus 5 above 22 percent utilisation and Kimi’s own API above 40 percent; the on-demand cluster at $287,000 beats Opus 5 above 46 percent and Kimi above 84 percent. Coding workloads follow the working day, so a cluster sized for the afternoon peak idles at night and at weekends. Forty percent is not a comfortable margin; 84 percent is not reachable. Figure 1 draws it.

