Built for the Real World · Essay · Open weights

Kimi K3: the weights you can hold, and the arithmetic before you do

On 27 July Moonshot AI published the weights of Kimi K3, 2.8 trillion parameters with 104 billion active, under a five-clause licence that asks nothing of internal users. It is the first open-weight model at the frontier, and it came with a price list built around a $0.30 cache hit. Whether an enterprise should hold the weights is a sum, shown here with every assumption labelled.

Abstract illustration for “Kimi K3: the weights you can hold, and the arithmetic before you do”

On Monday 27 July, Moonshot AI published the weights of Kimi K3 on Hugging Face, ten days after announcing the model on 17 July and on the day it had promised. The checkpoint is about 1.56 terabytes. The model is 2.8 trillion parameters, 104 billion of them active for any one token, with a one-million-token context and native vision. It ships under a five-clause document called the Kimi K3 License, which permits commercial use, asks nothing of anyone who runs it internally, and asks large hosting businesses for a separate agreement. By the evening, OpenRouter listed it from seven providers at the $3 per million input tokens and $15 per million output tokens that Moonshot charges itself. Moonshot’s price list carries a third number that matters more than either: $0.30 per million for input served from cache.

By its publisher’s benchmarks this is the first open-weight model at the frontier, and it arrived with a price list built around the cache rather than the token. For an enterprise the question is not whether the model is good. It is whether to hold the weights, and that depends on arithmetic almost no team has done: what it costs to keep 2.8 trillion parameters warm, what utilisation you need before that beats a $0.30 cache hit, and what you inherit when there is no vendor between you and the model. The arithmetic is below, every assumption labelled, and it is less flattering to self-hosting than the word “open” suggests.

What shipped, and the number that matters

In plain terms, K3 is a mixture-of-experts transformer. Each layer holds 896 expert networks, a router picks 16 of them for every token, and two shared experts see every token. That is how a 2.8-trillion-parameter model does the work of one a twenty-seventh of its size on each token: 104 billion parameters are active and the rest wait in memory. The model is 93 layers deep. Sixty-nine use Kimi Delta Attention, Moonshot’s linear-attention design, which keeps a fixed-size state rather than a key-value cache that grows with every token; 24 use Gated MLA, a gated form of the latent attention that compresses the key-value cache. That ratio is what makes a million-token context affordable to serve. Attention Residuals, the third name on the list, changes how information passes between layers; the technical report has the detail. Vision is native, through a 401-million-parameter encoder for images and video.

The precision is the part to read twice. Weights are MXFP4, a four-bit floating-point format with per-block scaling; activations are MXFP8; and Moonshot did quantisation-aware training from the supervised fine-tuning stage onward, so the four-bit weights are the model, not a lossy export of it. Two things follow. First, 2.8 trillion parameters at half a byte each is about 1.4 terabytes of weight data; with embeddings, the vision encoder and scaling factors, the checkpoint is about 1.56 terabytes. Second, MXFP4 runs natively on NVIDIA Blackwell and AMD MI400; on the Hopper generation most enterprises own, the weights would be unpacked into a wider format, roughly doubling the memory.

Why the active count matters for serving: compute per token scales with the 104 billion active parameters, memory with the 2.8 trillion total. Per token, K3 costs about what a 100-billion-parameter dense model costs to run, which is why Moonshot can sell output at $15 per million. Per deployment, it costs what a 2.8-trillion-parameter model costs to hold, which is why nobody runs it on one machine. The total sets the floor on hardware; the active count sets what you get once you have paid for it.

Moonshot reports 88.3 on Terminal-Bench 2.1, 67.5 on DeepSWE, 93.5 on GPQA Diamond and 91.2 on BrowseComp at maximum reasoning effort. Those are the publisher’s numbers, and the bench in the evaluation post exists because publisher numbers begin a qualification rather than end it.

The licence, clause by clause

The Kimi K3 License is 488 words long, a custom document rather than Apache or MIT, and custom documents are where open-weight projects stall. Here it is clause by clause, to be read with counsel.

The grant. MIT in form: permission, free of charge, to use, copy, modify, merge, publish, distribute, sublicense and sell the software, which includes weights, code and documentation, and to run, deploy, fine-tune and create derivative works. No payment appears anywhere.

Clause 1. The notice travels with all copies, and use must comply with applicable law. For most users that is the entire obligation.

Clause 2. The commercial gate, with two conditions that must both hold. Model as a Service means giving a third party access to inference or fine-tuning with meaningful control over inputs, parameters or training data; products with the model embedded in a feature, and the relaying of requests to a model hosted elsewhere, are excluded. If the licensee or any of its affiliates operates such a business, and the aggregate revenue of the licensee and its affiliates exceeds $20 million over any consecutive twelve months, the licensee must enter a separate agreement with Moonshot before using the software or its derivatives for any commercial purpose. Note what the revenue test measures: not the hosting business, but the group. A conglomerate with a cloud subsidiary selling hosted inference anywhere in its structure meets both conditions at any revenue a listed company has. The consequence is not a fee but a negotiation that must precede any commercial use.

Clause 3. A commercial product with more than 100 million monthly active users, or more than $20 million in monthly revenue, must display “Kimi K3” prominently on its interface. No fee, but for a bank or a telecom whose app clears either threshold, a model name on the screen is a brand decision and possibly a regulatory one.

Clause 4. The exemptions. Clauses 2 and 3 do not apply to internal use, defined as use that “does not make the Software, its outputs, or its underlying capabilities available to third parties”, nor to use through Moonshot’s own products or certified inference partners. That definition decides most enterprise cases. A coding assistant for your engineers, a document pipeline, a back-office classifier, an assistant for your own contact-centre staff: none makes the model or its outputs available to third parties, so they owe nothing and display nothing. A customer-facing chatbot does, so it is not internal use; it then meets clause 3 only above the thresholds and clause 2 only if the group runs a hosting business.

Clause 5. The software, and expressly its outputs, come as is, without warranty or liability. A hosted provider’s enterprise contract usually carries an SLA and some indemnity; here there is neither.

What the document omits is as informative. No governing law, so a dispute falls to whatever conflict-of-laws rules apply, with a licensor in Beijing. No patent grant, which Apache 2.0 has; no acceptable-use policy, no field-of-use restriction, no termination clause, no obligation to share improvements and no telemetry. Simon Willison noted on the day that Moonshot says “open weight” and never “open source”, and that is the right reading; the July comparison is Thinking Machines’ Inkling, 975 billion parameters released on 15 July under Apache 2.0, which a legal team clears without reading. For an Indian enterprise K3’s terms come down to three questions: is the use internal, does any affiliate sell hosted inference, and will the product cross 100 million users or $20 million a month? The answers belong in the model registry beside the router configuration described in the routing post, so that no workload can be pointed at the model under terms that exclude it.

What it costs to hold 2.8 trillion parameters

Memory first. An HGX B200 node carries eight GPUs with 180 gigabytes each, 1.4 terabytes in total; the checkpoint is 1.56 terabytes, so one node does not hold the model. Two nodes, sixteen GPUs and 2.88 terabytes, hold the weights with about 1.3 terabytes left for key-value cache and activations. That is the floor that fits, not the floor that performs: each GPU holds 56 experts of every layer, the router’s all-to-all traffic crosses the inter-node link on every token, and a million-token context for a few dozen concurrent users eats the spare memory quickly. Moonshot recommends a supernode of 64 or more accelerators in one interconnect domain: a GB200 NVL72 rack, 72 GPUs and 13.4 terabytes, or eight HGX nodes with expert traffic over InfiniBand. The smallest community quantisation, Unsloth’s one-bit dynamic GGUF, is 594 gigabytes, needs about 610 gigabytes of combined memory, agrees with the full model on 78.9 percent of top-1 predictions, and runs at roughly 20 tokens per second on a B200: usable for a lab, not for two hundred engineers.

Then the rent. By a July survey of rental rates, B200s run between $4.99 and $6.04 per GPU-hour on demand at specialist clouds, 36-month reserved contracts go as low as $2.25, and a GB200 NVL72 rack prices between $756 and $1,944 an hour. At $5.50 and 730 hours a month, the sixteen-GPU floor is about $64,000 a month and the 64-GPU floor about $257,000; reserved for three years at $2.25, the 64-GPU cluster is about $105,000. The rack is $552,000 to $1.42 million a month.

Then the people. Serving this model means expert-parallel inference across nodes, kernels for Kimi Delta Attention, an MXFP4 path that preserves quality, a prefix cache tier, observability, patching the stack every few weeks and on-call for a tier-one dependency. Moonshot publishes recipes for vLLM, SGLang and its own TokenSpeed; someone on your payroll still has to own them. I budget $30,000 a month for a three-person serving team with on-call cover: my estimate at Bengaluru rates, and the cheapest line in the sum.

Now the break-even. Self-hosting is a fixed cost, hardware plus team; the API is a variable cost, tokens times a blended price. The crossover is fixed cost divided by blended price, in tokens, and the utilisation you need is that count divided by the cluster’s capacity. The fixed cost is above. The blended price comes from the worked example below: at a coding agent’s token mix, with 90 percent of input from cache, Kimi’s API costs about $0.85 per million workload tokens and Claude Opus 5’s about $1.54. The capacity is the number nobody has published. I assume a 64-GPU cluster serves 400 billion workload tokens a month at full load in that mix, conservative since nine tokens in ten are cache hits that never touch a tensor core, and an estimate. On those assumptions the reserved cluster at $135,000 a month beats Opus 5 above 22 percent utilisation and Kimi’s own API above 40 percent; the on-demand cluster at $287,000 beats Opus 5 above 46 percent and Kimi above 84 percent. Coding workloads follow the working day, so a cluster sized for the afternoon peak idles at night and at weekends. Forty percent is not a comfortable margin; 84 percent is not reachable. Figure 1 draws it.

Monthly cost against utilisation: Kimi K3 API, Claude Opus 5 API, and self-hosted Kimi K3A line chart. Two rising lines show API cost growing with utilisation, Opus 5 steeper than Kimi K3. Two flat lines show self-hosted cost, on demand at $287,000 a month and reserved at $135,000. Clay markers show the break-even points at 22, 40, 46 and 84 percent utilisation. A panel on the right lists the assumptions. Monthly cost against utilisation of a 64-GPU B200 cluster serving Kimi K3 $0$150k$300k$450k$600k 0%25%50%75%100% Utilisation, as a share of the assumed capacity Claude Opus 5 API Kimi K3 API Self-hosted, on demand: $287k a month Self-hosted, 36-month reserved: $135k a month 22%40%46%84% ASSUMPTIONS Cluster: 64 B200, eight HGX nodes. On demand $5.50 per GPU-hour; 36-month reserved $2.25 (July 2026). 730 hours a month. Team $30k a month for three engineers (estimate). Capacity at 100%: 400 billion workload tokens a month. Estimate; no self-hosting throughput is published. Per call: 25,000 input tokens, 90% cache hits, 500 output (estimates). Kimi K3: $0.30 hit, $3 miss, $15 out, $0.85 per million blended. Opus 5: $0.50 hit, $6.25 cache write, $25 out, $1.54 per million blended. Prices verified 28 July 2026.
Figure 1. Self-hosting is a flat line and the API is a slope; the question is where they cross. At three-year reserved rates a 64-GPU cluster beats Opus 5 above about 22 percent utilisation and Kimi’s own API above about 40 percent. On demand, it never comfortably beats Kimi. The capacity assumption is the number to replace with your own measurement. Prices from Moonshot, Anthropic and a July 2026 GPU rental survey; everything else is my estimate.

The cache is the business model

Read Moonshot’s rate card as a business model. A cache miss on input is $3.00 per million tokens; a hit is $0.30; output is $15. Writing a prefix into the cache for five minutes costs $3.00, the miss price, so caching is free at the margin; an hour costs $6.00; every hit refreshes the entry without a further write charge. In four numbers the company has said it expects to serve most input from memory rather than compute. Mooncake, the architecture Moonshot described in a 2024 paper, separates prefill from decode into different pools of machines and keeps the key-value cache in a shared store across the cluster’s DRAM and SSD, so a prefix computed once is reused by any later request that shares it. On coding workloads Moonshot reports a hit rate above 90 percent.

That number is the economics of agents. As I argued in the July post on tokens to done, an agent re-sends its whole history at every step, and most of that history is identical to the previous step’s. A coding agent carrying 25,000 tokens of context pays $0.075 of input per step on a miss and $0.0142 at a 90 percent hit rate, a five-fold fall. Output does not cache, so a 500-token step costs $0.0075 regardless, and at that hit rate output is the largest line. The same step on Claude Opus 5, launched on 24 July at $5 input, $25 output, $0.50 for a hit and $6.25 to write a five-minute entry, costs about $0.027 of input and $0.0125 of output, roughly 1.8 times Kimi’s price. The gap between the APIs is a multiple. The gap between a cached and an uncached loop, on either, is the order of magnitude.

Two consequences. Prompt structure now has a dollar value: a stable prefix, an append-only history and unchanging tool definitions are worth a five-fold cut in input cost, and a timestamp at the top of the prompt throws it away silently. And the hosted provider’s cache is a shared, cluster-wide, engineered tier that makes the $0.30 price possible; self-host K3 with only per-node prefix caching and every step that misses the local node pays full prefill compute. A Mooncake-class cache tier is in Table 1 for that reason.

Sovereignty and control: what weights give you, and what you inherit

Weights give an enterprise five things an API cannot. Data that never leaves: prompts, outputs and the key-value cache stay on hardware you control, which satisfies the localisation rules Indian financial regulators impose on payment data, and the Digital Personal Data Protection regime’s concerns, by construction rather than by contract. Permanence: a hosted model can be retired, repriced or quietly retrained; the checkpoint you downloaded on 27 July is the checkpoint forever. Inspection: your own evaluation harness, refusal tests and adversarial suite, at no marginal price. Modification: fine-tuning on proprietary data, distillation into smaller task models, changes to the inference stack, all permitted. And independence from a rate limit, a terms-of-service change or a foreign policy decision. India has provisioned more than 38,000 GPUs through the IndiaAI Mission’s compute portal, for startups and academia rather than enterprises; the national supply of accelerators is no longer the binding constraint on a sovereign deployment. The arithmetic is.

What you inherit is the rest of the job. A hosted frontier provider runs classifiers on inputs and outputs, monitors abuse across its whole customer base, maintains an incident process and ships refusal training tested at scale. Run K3 yourself and you get whatever refusal behaviour Moonshot trained in, with no gate in front of it. You build the filtering, the egress controls for any agent that holds a credential, the logging, the red-team cycle and a stop path someone has tested. You own uptime and on-call, the checksums and provenance of a 1.56-terabyte artefact, and licence compliance across every subsidiary that touches the model. And the model never improves: when Moonshot ships a better checkpoint, you download, re-qualify and redeploy, and the qualification is the slow part. A hosted provider stands partly behind its model. Here, you stand entirely behind yours.

None of this argues against holding the weights. It argues for pricing the walls before the racks. Table 1 puts the licence conditions beside the operational prerequisites, because both lists go to the same approval meeting.

ItemWhat the licence says, or what self-hosting requiresWhat it means for an enterprise
Grant
Licence · preamble
Use, copy, modify, distribute, sublicense, sell, run, deploy, fine-tune and make derivative works, free of charge, subject to the clauses belowMIT-form rights to weights and code. No payment anywhere in the document
Notice and law
Licence · clause 1
Keep the copyright and permission notice with all copies; comply with applicable lawShip the notice with every deployment artefact. The whole obligation for internal users
Hosting gate
Licence · clause 2
If the licensee or any affiliate runs a Model as a Service business, and group revenue exceeds $20M over any 12 months, a separate agreement with Moonshot is needed before any commercial useRead “affiliates” twice. A group with a hosted-inference subsidiary is caught at any revenue a listed company has. Embedded features and request relaying are excluded
Attribution
Licence · clause 3
Commercial products above 100M monthly active users or $20M monthly revenue must display “Kimi K3” prominently on the interfaceA brand and, in regulated sectors, a disclosure decision for large consumer products. No fee
Exemptions
Licence · clause 4
Clauses 2 and 3 do not apply to internal use (model, outputs and capabilities not made available to third parties) or to access via Moonshot’s products and certified partnersCoding assistants, back-office pipelines and staff-facing tools owe nothing and display nothing. Customer-facing products are not internal use
Warranty
Licence · clause 5
Software and its outputs provided as is; no warranty, no liabilityNo indemnity and no SLA. Your contract with your users is the only one that exists
Absent
Licence · not stated
No governing law, no patent grant, no acceptable-use policy, no termination clause, no telemetry, no obligation to share improvementsFewer restrictions than most custom model licences, and less certainty about how a dispute would be resolved
Memory
Self-hosting
1.56 TB checkpoint. One HGX B200 node (1.4 TB) does not hold it; 16 B200s (2.88 TB) fit with ~1.3 TB spare; Moonshot recommends 64+ accelerators in one supernodeTwo nodes is the floor that fits; eight nodes or an NVL72 rack is the floor that performs. Hopper needs wider formats and double the memory
Rent
Self-hosting
B200 $4.99 to $6.04 per GPU-hour on demand, $2.25 reserved for 36 months; GB200 NVL72 rack $756 to $1,944 an hour (July 2026 survey)$64k a month for 16 GPUs; $257k on demand or $105k reserved for 64; $552k to $1.42M for a rack
Serving stack
Self-hosting
Expert-parallel inference across nodes, Kimi Delta Attention kernels, MXFP4 path; Moonshot recipes for vLLM, SGLang and TokenSpeedSomeone on your payroll owns it, patches it and is on call for it
Cache tier
Self-hosting
Cluster-wide prefix cache in DRAM and SSD, Mooncake-style, to reach the 90% hit rate Moonshot reports on codingWithout it, every step pays full prefill compute and the break-even in Figure 1 moves right
Team
Self-hosting
Three engineers with on-call cover; my estimate $30k a month at Bengaluru ratesThe cheapest line in the sum, and the one most often left out of it
Containment
Self-hosting
Input and output filtering, egress control for agents, logging, red-team cycle, tested stop path, checksum and provenance of the weightsNone of it ships with the model. The hosted provider’s controls are the part of the API price you do not see
Qualification
Self-hosting
Run the model on your own bench before and after every checkpoint change; record the licence classification in the model registryPublisher benchmarks start the process. The bench and the registry finish it
Table 1. The Kimi K3 License, clause by clause, and the operational prerequisites of running the model yourself. Both lists go to the same approval meeting. Licence text from the Hugging Face repository, 27 July 2026; hardware figures from NVIDIA’s HGX B200 specification and a July 2026 rental survey.

A worked example: 200 engineers, three ways

Take a coding-assistant workload for 200 engineers, the deployment most groups consider first. Every token count below is an estimate; the prices are verified as of today. Assume each engineer runs 20 agent tasks a working day, each task runs 20 model calls, and the month has 22 working days: 1.76 million calls. Assume each call carries 25,000 tokens of context, 90 percent served from cache, matching Moonshot’s reported rate on coding work and applied to both APIs, and produces 500 output tokens. That is 44 billion input tokens a month, 39.6 billion cached and 4.4 billion not, and 0.88 billion output tokens.

Kimi K3 through Moonshot’s API. Cached input, 39.6 billion tokens at $0.30 per million, is $11,880. Uncached input, 4.4 billion at $3.00, is $13,200. Output, 0.88 billion at $15, is $13,200. About $38,300 a month, or $191 per engineer: three nearly equal lines, the signature of a well-cached agent.

Claude Opus 5 through Anthropic’s API. Cached input at $0.50 is $19,800. Uncached input, billed as five-minute cache writes at $6.25 because an agent loop writes every new prefix, is $27,500; at the base rate of $5 it would be $22,000. Output at $25 is $22,000. About $69,300 a month, or $346 per engineer. The cache write is the line people forget, and here it is the largest.

Kimi K3 self-hosted on rented GPUs. The sixteen-GPU floor that fits, at $5.50 per GPU-hour, is $64,240 a month plus the $30,000 team: about $94,000, and it would not serve 200 engineers at acceptable latency. The 64-GPU floor that performs is $257,000 on demand or $105,000 reserved, plus the team: about $287,000 or $135,000. Neither carries a cache tier, a containment stack or any discount for nights and weekends.

At 200 engineers, then, Moonshot’s API is cheapest by a wide margin, Anthropic’s is 1.8 times the price, and self-hosting is 2.5 to 7.5 times the cheapest API depending on hardware and commitment. The self-hosted line crosses Kimi’s API only when the workload is about 3.5 times this size at reserved rates, roughly 700 engineers’ worth of agent traffic on one cluster, and crosses Opus 5 at about twice this size. Below those multiples, an enterprise holds the weights because of a constraint on the data, not a saving on the bill.

What to do

  1. Do the arithmetic with your own numbers before you buy anything. Run the workload on the API for a month and measure tokens per task, cache-hit rate and the daily shape of demand. Two of the break-even formula’s three inputs only you can supply.
  2. Classify the licence once and record it. Internal use, customer-facing below the thresholds, customer-facing above them, or a group with a hosting affiliate: four classes, four obligations. Put the class in the model registry beside the router configuration.
  3. Hold the weights as an option, not a deployment. Download, verify and archive the checkpoint, and keep a tested recipe for bringing it up on rented capacity, so that when a data constraint or a provider decision closes the hosted route, exercising the option takes days, not quarters.
  4. If you self-host, size for Moonshot’s floor and budget for the cache tier. Two nodes fit the model and will disappoint everyone who uses it. Reserve capacity only after a measured month shows the utilisation that justifies it, and treat the cluster-wide prefix cache as part of the system, because the price you are trying to beat is the price the cache made possible.
  5. Price the walls before the racks. Filtering, egress control, logging, red-teaming and a tested stop path are the part of a hosted provider’s price you do not see. If that budget is not approved, the hosted route with the provider’s controls is the safer one even where it is the dearer one.

On Sunday the frontier was closed. On Monday a model at or near it can be downloaded by anyone with 1.56 terabytes of disk, under terms that ask nothing of most of the people who will run it. That is a genuine change, and the people who built it priced it with unusual honesty: the cheapest way to use the open model is still to let its makers run it, because the serving system around the weights is where the economics live. The weights you can hold are worth holding. Whether they are worth running is a sum, and for most enterprises in July 2026 the sum says not yet.

Sources

  1. Moonshot AI. Kimi K3. 17 July 2026.
  2. Moonshot AI, on Hugging Face. Kimi K3 License. 27 July 2026.
  3. Moonshot AI, on Hugging Face. moonshotai/Kimi-K3 model card. 27 July 2026.
  4. Moonshot AI. Kimi Open Platform: chat model pricing. Retrieved 28 July 2026.
  5. Simon Willison. Kimi K3. 27 July 2026.
  6. Unsloth. Kimi K3: how to run locally. July 2026.
  7. Anthropic. Claude API pricing. Retrieved 28 July 2026.
  8. IntuitionLabs. Data center GPU pricing 2026. Updated 20 July 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn