On Wednesday 12 August, Alibaba published the weights of Qwen3.8-Max, a 2.4-trillion-parameter mixture-of-experts model with 95 billion parameters active per token and a one-million-token context, nine days after making it available through its own cloud. The day before, NVIDIA released Nemotron 3.5 Lightning, a 31.6-billion-parameter hybrid Mamba-Transformer with 3.6 billion active, under a licence that permits commercial use without material restriction. Two weeks before that, Moonshot had published the weights of Kimi K3, 2.8 trillion parameters, and a month before that Thinking Machines had released Inkling, 975 billion parameters, under Apache 2.0. Yesterday Google shipped Gemini 3.7 Flash at half the introductory price of 3.6 Flash, which is closed, and which is the model all of these will be compared to.
Four open-weight models at or near frontier scale in five weeks. Two of them are larger than anything the Western labs have released openly. One of them was trained in the United States by people who left a frontier lab to build it. For an enterprise, and for a country, this changes a question that has been theoretical for two years. Sovereignty over AI, in the sense of running a frontier-class model on infrastructure you control, with data that never leaves it, under terms you can read, is now a procurement option rather than a policy aspiration. The question is what it actually costs, and what it does not buy.
What is on the table
Qwen3.8-Max. Alibaba’s flagship, released through QwenCloud on 3 August at 10 percent of its eventual price during preview, with weights following on 12 August as Qwen3.8-2.4T-A95B. On Alibaba’s own benchmarks it scores 86.6 on TerminalBench 2.1 against GPT-5.6 Sol’s 88.8, 93 on PaperBench, and “near or above” Claude Opus 4.8, Claude Fable 5 and GPT-5.6 Sol across many categories. The long-horizon demonstrations are the striking part: a 16-day autonomous software project with 265 commits, a five-day reproduction and improvement of a research paper, a full fiscal year of a simulated e-commerce business, five hundred iterations of chip-design optimisation. It speaks both the OpenAI and Anthropic API protocols. What it does not have, as of today, is a published licence. The weights are downloadable; the terms under which you may use them commercially are not stated anywhere in the release.
Kimi K3. Moonshot’s 2.8-trillion-parameter model, 16 of 896 experts active per token, one-million-token context, shipped in MXFP4 weights with MXFP8 activations. Weights are about 1.4 terabytes before runtime overhead; Moonshot recommends a “supernode” of 64 or more accelerators, and no single current GPU holds it. It ranks fourth on the Artificial Analysis Intelligence Index at 57 and second on the Vals Index. The licence is described by several sources as a modified MIT with a revenue-share clause of up to 30 percent for inference providers earning more than $20 million a year from it; Moonshot’s own licence text was not published at announcement, and the clause should be read before it is relied on.
Inkling. Thinking Machines’ first model, 975 billion parameters, Apache 2.0, released on 15 July. Apache is the licence enterprises know how to approve, and that, more than any benchmark, is why Inkling matters: it is the first frontier-scale model from a US lab that a corporate legal team can clear in an afternoon.
Nemotron 3.5 Lightning. The opposite end of the scale and, for most enterprise workloads, the more useful release. 31.6 billion parameters total, 3.6 billion active, a hybrid Mamba-Transformer design, NVFP4 quantisation alongside BF16, a one-million-token context, and the OpenMDW-1.1 licence, which publishes weights, training recipe and data details and permits commercial modification. It scores 24 on the Intelligence Index, level with gpt-oss-120b at a quarter of the parameters, and produces nearly 670 tokens per second on a hosted endpoint. NVIDIA positions it as the execution layer for long-running agents: the model that does the thousand cheap steps while a frontier model does the ten expensive ones.
What “sovereign” actually requires
Sovereignty is used loosely. For an enterprise it decomposes into five things, and an open-weight model delivers exactly one of them.
Weights. You hold the model file. This is what the releases above give you, and it is necessary: without it, every other control depends on a vendor. It is also the easy part.
Licence. Terms that permit your use. Apache and OpenMDW clear a legal review quickly. A revenue-share clause triggers a commercial negotiation. An absent licence, which is Qwen3.8-Max’s position today, blocks production use at any organisation with a functioning legal department, however good the model is. The gap between “weights available” and “weights usable” is where most sovereignty projects stall.
Compute. The capacity to run it. A 2.4- or 2.8-trillion-parameter model needs a cluster that most enterprises do not own and most national programmes are only beginning to provision. India’s IndiaAI Mission has sanctioned 93 lakh GPU-hours and backed twenty indigenous model proposals, five of them released; Sarvam has trained a 30-billion and a 105-billion-parameter model on that compute. That is the scale at which a sovereign programme can currently train. Running a trillion-scale open model is a different, smaller problem, but it is still a problem measured in racks.
Data. Inference on infrastructure where the data stays. This is the requirement regulated industries actually have, and open weights satisfy it trivially while hosted APIs satisfy it only when the provider builds an in-country region, which is why the hosted labs are racing to do so.
Controls. Safety, monitoring, refusal behaviour, incident response. This is what open weights do not come with, and the summer’s containment incidents are the case study in what it costs to lack them. When you run Qwen3.8-Max yourself you inherit its full capability, including the capability a hosted provider would have gated, and you inherit the whole job of containing it.
The India view
I have spent most of my career building language technology for Indian languages and Indian enterprises, including a year at Krutrim working on Indic foundation models, so I read this wave with a particular interest. India has a national programme, the IndiaAI Mission, that has provisioned compute, backed twenty sovereign model proposals and seen five released, with Sarvam’s 30-billion and 105-billion-parameter models the most visible. Sarvam became a unicorn in June on a $234 million round led by HCLTech, and on 30 July it moved up the stack with Indus Work Agents, a coding assistant and a voice-agent platform opened to all developers, reporting more than 500 enterprise and startup customers. Krutrim, by contrast, announced in May that it was halting its own chip and model work to focus on cloud services, and posted a profit doing so. Two Indian frontier efforts, two opposite answers to the question of whether to build the model or build on one.
The open-weight wave makes the second answer stronger. A 2.4-trillion-parameter model trained by someone else, run on compute you control, with Indian-language data that never leaves the country, is a sovereign deployment in every sense that matters to a regulator, and it is available now, without a training programme. What it does not give you is the Indic-language depth that a model trained for it has, nor the controls, nor, in Qwen’s case today, a licence. The practical sovereign stack for an Indian enterprise in August 2026 is therefore layered: an Indic-tuned model where language quality is the product, an open-weight generalist where data residency is the constraint, and hosted frontier models, in-country where offered, where capability is the constraint and the data permits it.
Where open weights win, and where they do not
They win on bounded, high-volume work. Nemotron 3.5 Lightning, at 3.6 billion active parameters and 670 tokens per second, on a single GPU, under a permissive licence, is the right model for the classification, extraction and routing steps that dominate agent loops, and it is a model you can run in an air-gapped facility. This is the open-weight release of the summer that will actually change enterprise bills.
They win on data that cannot leave. Regulated records, source code under export control, material non-public information. Where the data residency requirement is absolute, the hosted option does not exist unless the provider has built a region, and even then the contract is the control.
They win on customisation. Fine-tuning on proprietary data, distillation into smaller task models, modification of the inference stack. OpenMDW’s publication of recipe and data is the first licence that makes this a documented path rather than reverse engineering.
They lose on operations. Running a trillion-parameter model is a platform-engineering programme: multi-node serving, quantisation, caching, observability, patching the serving stack, and on-call. The hosted labs amortise that across millions of customers. A single enterprise amortises it across itself.
They lose on controls. The summer’s incidents happened inside the labs that wrote the safety frameworks. An enterprise running an unguarded 2.4-trillion-parameter model inherits the full containment problem with none of the lab’s monitoring, refusal training or incident apparatus. The weights are free. The walls are not.
They lose, for now, on the frontier. Alibaba’s own table puts Qwen3.8-Max two points behind GPT-5.6 Sol on TerminalBench, and Sol is not the frontier any more. The gap is months, not years, and it is closing, but it is real, and for the small set of workloads where the last few points matter it decides the choice.
The real bill for running a trillion-parameter model
“Open weights are free” is true in the sense that a free puppy is free. Here is the arithmetic for Kimi K3, which is the one that has published enough to estimate.
The weights are about 1.4 terabytes in MXFP4. Moonshot recommends a supernode of 64 or more accelerators, because no single current GPU holds the model and mixture-of-experts routing across nodes needs fast interconnect. That is, at the moment, a rack-scale system from one of three vendors, with a list price in the low millions of dollars and a lead time measured in quarters, or a reserved cluster from a cloud provider at a monthly rate that competes with a hosted frontier subscription for all but the heaviest users. On top of hardware: a serving stack that handles expert parallelism, a quantisation path that preserves quality, a prompt cache, observability, patching, and an on-call rotation for a system that is now a tier-one dependency. Our estimate for a single-tenant deployment at production scale is two to three platform engineers and a six-figure monthly infrastructure line before a single token is served.
Against that, the hosted price for the same capability class is a few dollars per million tokens with none of the operations. The crossover depends on volume and on how much of the hosted price is compute versus margin, and for most enterprises it does not arrive until the workload is very large or the data constraint makes the comparison irrelevant. Nemotron 3.5 Lightning is a different sum entirely: a 25-gigabyte quantised model that runs on a single GPU, or on a high-end consumer card, with a published recipe and a permissive licence. The bill for that is an afternoon and a machine you probably already have. The open-weight wave is really two waves, one that changes what is possible for a national programme and one that changes what is cheap for a platform team, and they should be evaluated separately.
How to read an open-weight licence
Because the licence is where sovereignty projects stall, a short guide. Four questions decide whether a model can go into production at an organisation with auditors.
Is there a licence at all? Qwen3.8-Max, as of today, has weights and no terms. Downloading is possible; use is undefined. That is not a technicality. Without terms there is no grant of rights, and a legal department will treat the model as unlicensed software, which is to say unusable.
Is it a known licence or a custom one? Apache 2.0, which Inkling uses, is understood by every legal team and carries a patent grant. OpenMDW-1.1, which Nemotron uses, is newer and specific to models but is drafted to be permissive about commercial use and modification and publishes recipe and data alongside weights. A custom licence, which is what Kimi K3 reportedly has, needs to be read in full, and the reading takes time.
Does it have a commercial trigger? Kimi K3’s reported revenue-share clause applies to inference providers above $20 million in annual revenue from the model. For an enterprise using the model internally that may never bite; for one that resells access it will. Clauses like this are becoming common in Chinese open releases and they are enforceable, so the question is not whether to accept them but whether your use case crosses the trigger.
Does it restrict use? Acceptable-use policies attached to model licences can exclude whole industries or applications. Read the exclusions against your actual deployments, not your intended ones. A licence that permits everything except a field one of your subsidiaries operates in is a licence you cannot adopt group-wide.
The practical rule we have adopted is that a model enters the production catalogue only when its licence has been reviewed and classified, and that the classification travels with the model in the router’s configuration, so that a workload cannot be pointed at a model whose terms exclude it. This sounds bureaucratic until the first time a team deploys a model with an unpublished licence into a regulated process.
A posture, not a decision
- Classify workloads by what constrains them: data residency, capability, cost, language. Each constraint points at a different layer of the stack, and almost no organisation has one constraint everywhere.
- Adopt one permissive open-weight model for the bounded tier now. Nemotron 3.5 Lightning or an equivalent under Apache or OpenMDW, run on your own infrastructure, behind the same router as everything else. This is low risk, high volume and immediately cheaper.
- Treat trillion-scale open models as an option to exercise, not a default. Stand up the capability to run one, on a cluster you can provision, and benchmark it against the hosted frontier on your own tasks. Exercise the option when a workload’s data constraint makes the hosted route impossible, or when the licence and the benchmark both clear.
- Do not run a model without a licence. Qwen3.8-Max is the best open model released this year and, until Alibaba publishes terms, it is unusable in production at any organisation that answers to auditors.
- Budget for the walls. Sandboxing, egress control, monitoring from outside the model, and a tested stop path cost real engineering time, and they are the price of running an unguarded model. If that budget is not approved, the hosted route, with the lab’s controls, is the safer one even when it is the more expensive one.
Five weeks ago the frontier was closed and the open models were a tier below it. Today the gap is a few benchmark points and a licence file. That is a different world for anyone whose data cannot leave the building, and for any country that wants its AI to run at home. It is not, yet, a world where open weights replace the hosted labs. It is one where the enterprise finally has a real choice, and the choice has to be made per workload, with the walls priced in.