The architecture that won the argument: state space models, three years on
Mamba promised linear time, constant memory and no key-value cache, and nearly three years later no pure state space model runs at the frontier. Yet Kimi K3, Qwen3.8-Max and Nemotron 3.5 Lightning all keep attention in about one layer in four. The fixed state lost the ecosystem argument and won the architecture argument by stealth, as a layer type. Here is why, with the papers and the memory arithmetic.
On Monday 27 July, Moonshot AI published the weights of Kimi K3, and the technical report that came with them lists 93 layers. Sixty-nine of them are not attention. They are Kimi Delta Attention, a linear-attention layer that keeps a fixed-size state and no key-value cache; the other 24 are a gated form of the latent attention the rest of the industry would recognise. Sixteen days later Alibaba released Qwen3.8-Max as a 2.4-trillion-parameter checkpoint with 92 layers, 69 of them Gated DeltaNet and 23 of them attention.
The day before that, NVIDIA shipped Nemotron 3.5 Lightning with 52 layers, six of which attend. Three of the largest open-weight releases of the summer, from three companies on two continents, agree that about three layers in four of a frontier model should not be attention.
And yet. Mamba, the paper that argued a state space model could be the whole network, was posted on 1 December 2023, two years and ten months ago. The largest pure Mamba with published weights that I can find is 7 billion parameters. The company that went furthest with linear attention at scale, MiniMax, went back to full attention in October 2025 and wrote down why. Google’s Gemma 4 uses no linear layers at all. The cleanest story in sequence modelling, linear time, constant memory, no cache, did not displace attention.
It did something stranger and, for a platform team, more useful: it stopped being an architecture and became a layer type. This essay is about why, with the papers, and about what it changes when you buy and serve models. I have run inference fleets for three employers; the memory arithmetic at the end is the part I wish someone had shown me earlier.
The state space idea, from first principles
Start with a linear recurrence. A state space layer keeps a vector h, the state, and at each step updates it with the new input and reads an output from it: ht = A ht−1 + B xt, then yt = C ht. A decides what the state keeps from the previous step, B decides how the new token is written in, C decides what is read out.
In S4, the paper Albert Gu, Karan Goel and Christopher Ré posted on 31 October 2021 (arXiv 2111.00396), A, B and C are fixed for the whole sequence, and that one fact buys two things. Because the layer is linear and time-invariant, unrolling the recurrence gives a convolution: the output is the input convolved with a kernel whose entries are CB, CAB, CA²B and so on, which can be computed for a whole sequence at once with a fast Fourier transform.
So training is parallel, like a transformer, and inference is a recurrence, like an RNN, with a state that does not grow. S4’s contribution was a parameterisation of A that made that kernel computable stably. The result took every task on the Long Range Arena, including Path-X at 16,000 steps, which no previous model had solved, and generated about 60 times faster than a transformer of similar size.
What it could not do was language. Hungry Hungry Hippos, from Daniel Fu, Tri Dao and colleagues on 28 December 2022 (arXiv 2212.14052), found the reason with two synthetic tasks: recalling an earlier token, and comparing tokens across the sequence. A layer that treats every position identically cannot decide to remember this token and forget that one. H3’s answer was a layer built for those two operations, and a telling footnote: a 125-million-parameter model that kept two attention layers alongside H3 beat a transformer on OpenWebText by a full point of perplexity. The hybrid was in the literature a year before Mamba.
Mamba’s move, in Gu and Dao’s paper of 1 December 2023 (arXiv 2312.00752), was selection: let B, C and the step size depend on the current token. Now the layer can write a token into the state strongly or barely at all, and reset the state or carry it. On a selective copying task the full model scores 99.8 percent where a non-selective S4 layer scores 18.3. On induction heads it trains at length 256 and generalises to about a million tokens.
The price of selection is that the convolution view is gone, because the kernel is no longer the same at every position, so Mamba had to run as a recurrence during training. The paper’s second contribution is the hardware trick that made that fast enough: fuse the whole scan into one kernel, materialise the expanded state in the GPU’s on-chip SRAM rather than in high-bandwidth memory, use a parallel scan, and recompute states in the backward pass rather than store them.
A 3-billion-parameter Mamba matched transformers of twice its size on downstream tasks, and because there is no cache, inference throughput was four to five times a transformer’s, mostly because the batch can be larger.
Here is the thing to hold on to. In every one of these models the state has a fixed size, a small number of channels per feature dimension, 16 in Mamba, set at design time. Everything the model knows about the previous million tokens has to fit in it. A transformer keeps a key and a value for every token in every layer, which is the key-value cache, and that is why its memory grows with context and its per-token cost grows with context. The fixed state is what buys linear time and constant memory.
It is also a lossy compressor of the past, and every failure mode in the next section is what happens when you ask a compressor to do what an index does.
Figure 1. The trade in one picture. An attention layer stores a key and a value for every token and can read any of them back; its cache grows with the sequence. A selective state space layer folds each token into a fixed-size state, chosen per token, and its memory does not grow at all. After Gu and Dao, arXiv 2312.00752, December 2023.
Three ways a fixed state fails, with the papers
Copying. ‘Repeat After Me’, from Samy Jelassi, David Brandfonbrener, Sham Kakade and Eran Malach on 1 February 2024 (arXiv 2402.01032), is the cleanest statement. Theory first: a two-layer transformer can copy strings whose length is exponential in the number of heads, by attending back to the position it is copying from; any model with a fixed state, which the paper calls a generalised state space model, needs state that grows linearly with the string, because a state holding fewer than L log D bits cannot reproduce L tokens from a vocabulary of D.
Then the experiments, which are worse for the state space model than the theorem. On synthetic copying of 300-token strings, transformers needed about a hundred times fewer training examples. On a phone-book lookup, the smallest Pythia model, 410 million parameters, beat the largest Mamba, 2.8 billion, once the book had 70 entries. On question answering, the two 2.8-billion models did equally well on short passages, and Mamba’s score fell faster as the passages got longer.
Recall. The Zoology paper, from Simran Arora and colleagues at Stanford on 8 December 2023 (arXiv 2312.04927), asked why gated convolutions trailed attention by up to 2.1 points of perplexity on the Pile, and found that 82 percent of the gap sat in about 6 percent of tokens: the ones that require recalling something seen earlier in the context, a name, a number, the second half of a phrase. A 70-million-parameter attention model did better on those tokens than a 1.4-billion-parameter Hyena.
The paper formalised the task as multi-query associative recall and found that a two-layer gated convolution needs its model dimension to grow with the sequence length to solve it, where attention solves it at a fixed dimension of 64. The follow-up, Based, on 28 February 2024 (arXiv 2402.18668), turned that into a design rule: linear attention plus a short sliding window, with the state size as a dial you turn against memory, and 6.22 points more on recall-heavy tasks than Mamba at the same perplexity. The same group’s hybrids closed 97.4 percent of the gap with a few attention layers. That number is the whole hybrid story in advance.
State tracking. ‘The Illusion of State in State-Space Models’, from William Merrill, Jackson Petty and Ashish Sabharwal on 12 April 2024 (arXiv 2404.08819), proved that SSMs of the Mamba and S4 kind sit in the same complexity class as transformers, TC⁰, and so cannot compose permutations, which means they cannot reliably follow a chess game in notation, evaluate code that mutates variables, or track entities through a long narrative, despite having something called a state. The restriction turns out to be specific: Mamba’s transition values are confined to the range between zero and one.
Riccardo Grazzi and colleagues showed on 19 November 2024 (arXiv 2411.12537) that letting the eigenvalues go negative lets Mamba and DeltaNet learn parity and, with the right structure, any regular language. RWKV-7, posted on 18 March 2025 (arXiv 2503.14456), built its generalised delta rule around that and claims state tracking beyond what a transformer can do. Mamba-3, from Gu, Dao and six co-authors on 16 March 2026 (arXiv 2603.15569), adds a complex-valued state update for the same reason. So the third limit is being engineered away. The first two are not, because they are not limits of the transition. They are limits of the size of the state.
Read the three together and they are one fact. Copying needs memory linear in the length of what you copy. Recall needs the key-value pairs stored somewhere they can be looked up by content. State tracking needs the transition to be expressive. A fixed state can be made more expressive. It cannot be made bigger without giving up the reason it exists.
Linear time, constant memory, and still not the architecture of a single frontier model
The hardware story
Sara Hooker’s essay ‘The Hardware Lottery’, posted on 14 September 2020 (arXiv 2009.06489), argues that research ideas win or lose partly on whether they fit the hardware and software that happen to exist. The transformer has been winning that lottery since 2017, and the winnings compounded. FlashAttention, from Tri Dao and colleagues on 27 May 2022 (arXiv 2205.14135), made exact attention IO-aware, tiling the computation so that the attention matrix never touches high-bandwidth memory; it trained GPT-2 three times faster and produced the first transformer to beat chance on Path-X.
FlashAttention-2, on 17 July 2023, reached 225 teraflops per second on an A100, about 72 percent of the chip. FlashAttention-3, on 11 July 2024, reached 740 teraflops in FP16 on an H100, about 75 percent utilisation, and close to 1.2 petaflops in FP8. Every inference engine, every quantisation scheme, every serving trick, the paged cache, speculative decoding, prefix caching, was built on the assumption that the thing being cached is a key-value cache.
A selective scan is a worse fit for the same hardware, and the people who built it said so first. Mamba’s scan is a recurrence; it does not run on the matrix-multiplication units that make a modern GPU fast, and Mamba-2’s own paper notes that the scan ‘does not leverage matrix multiplication units’. The kernels were in Triton, which is productive and portable and carries launch overheads that bite at small batch sizes, which is precisely where interactive serving lives.
The state is sensitive to precision: Quamba, on 17 October 2024 (arXiv 2410.13229), found that the inputs to the scan carry outliers that attention does not, and MiniMax’s engineers reported in 2025 that linear attention tolerates the low-precision state storage inference engines rely on less well than attention does. None of this is a law of nature. All of it is a decade of engineering that one side had and the other did not.
Mamba-2 is the paper that tried to close the gap on the hardware’s terms. ‘Transformers are SSMs’, Dao and Gu, 31 May 2024 (arXiv 2405.21060), shows that a selective SSM with a scalar decay is equivalent to a structured matrix, a semiseparable one, and that the same matrix is a masked form of linear attention. That is the state space duality of the title, and its practical payload is an algorithm, SSD, that computes the scan in blocks using matrix multiplications.
The result was two to eight times faster than Mamba’s fused scan, overtook FlashAttention-2 at sequence length 2,000, and allowed a state eight times larger at almost no cost, which, by the recall argument above, is the dimension that matters. A 2.7-billion Mamba-2 trained on 300 billion tokens beat Mamba-2.8B and Pythia-6.9B.
The duality also meant that Mamba-2 and the linear-attention family, RetNet from Microsoft on 17 July 2023 (arXiv 2307.08621), DeltaNet on 10 June 2024 (arXiv 2406.06484), and Gated DeltaNet from Songlin Yang, Jan Kautz and Ali Hatamizadeh on 9 December 2024 (arXiv 2412.06464), which added Mamba’s gating to the delta rule and beat Mamba-2 on recall, were now one family with one set of kernels to optimise. Mamba-3’s multi-input, multi-output variant pushes further in the same direction, raising arithmetic intensity so the hardware does more per byte moved. The architecture had learned to speak matmul. It had not learned to retrieve.
What the large trainings showed, and the rise of the hybrid
The honest test of a pure SSM is a large one trained on the same data as a transformer, and NVIDIA ran it. Roger Waleffe and colleagues, with Gu and Dao as co-authors, posted ‘An Empirical Study of Mamba-based Language Models’ on 12 June 2024 (arXiv 2406.07887): 8-billion-parameter Mamba, Mamba-2 and transformer models trained on up to 3.5 trillion tokens. The pure models matched or beat the transformer on most tasks and fell behind on the ones the theory predicted, five-shot MMLU, a phone-book task, long-context reasoning.
The hybrid, 43 percent Mamba-2 layers, 7 percent attention and 50 percent MLP, beat the transformer by 2.65 points on average across twelve tasks and was projected to generate up to eight times faster. The pure SSMs that followed stopped at 7 billion: Codestral Mamba from Mistral on 16 July 2024, and Falcon Mamba from TII, trained on 5.8 trillion tokens and posted on 7 October 2024 (arXiv 2410.05355) as the first competitive attention-free model at that size. Competitive at 7 billion is a real result. Nobody published a pure one at 70.
AI21’s Jamba, 28 March 2024 (arXiv 2403.19887), was the first production hybrid, and its ablations are still the clearest statement of why. One attention layer for every seven Mamba layers, four of 32, 52 billion parameters with 12 billion active, a 256,000-token context, and a cache of about 4 gigabytes at that length against 32 for Mixtral. At 1.3 billion parameters a 1:3 ratio and a 1:7 ratio performed the same, so they took the cheaper one.
And the pure Mamba they trained alongside failed in a way no benchmark average shows: on IMDB sentiment it scored 48.8 against 84.1 for attention, because asked to answer ‘Positive’ or ‘Negative’ it would answer ‘Very Good’ or ‘3/10’. It had the knowledge and could not copy the format from the prompt. Add four attention layers and the score was 90.9.
Mamba-2’s authors found the same shape in their own ablation: in a 48-layer model, perplexity improved as attention blocks were added, was best at about six, and worsened beyond that; their hypothesis was that SSM layers do the general sequence work and attention layers act as a retrieval mechanism.
What followed is in Table 1, and the pattern is the finding. Between March 2024 and August 2026 every organisation that trained a large model with state space or linear layers kept some attention, and the fraction it kept fell into a narrow band.
Model
Linear layer
Linear : attention
Size
What was reported
Jamba AI21 · 28 Mar 2024
Mamba-1
7 : 1 (4 of 32)
52B · 12B active
256K context; ~4 GB cache at 256K vs 32 GB for Mixtral; pure Mamba failed format following (IMDB 48.8 vs 84.1)
Zamba Zyphra · 26 May 2024
Mamba-1
one shared attention block every 6
7B · 1T tokens
Strongest non-transformer at the size; shared block cuts attention parameters and cache
Mamba-2-Hybrid NVIDIA · 12 Jun 2024
Mamba-2
43% Mamba-2, 7% attention, 50% MLP
8B · 3.5T tokens
+2.65 points over an 8B transformer on 12 tasks; pure Mamba lagged on MMLU and phone-book
MiniMax-01 MiniMax · 14 Jan 2025
Lightning attention
7 : 1 (80 layers)
456B · 45.9B active
1M training context, 4M at inference; pure linear attention had ‘limited retrieval capabilities’
Nemotron-H NVIDIA · 21 Mar 2025
Mamba-2
10 attention of 118 (56B); 4 of 52 (8B)
56B · 20T tokens
Up to 3× faster inference than Qwen-2.5-72B and Llama-3.1-70B at 65K-token inputs
Falcon-H1 TII · 20 May 2025
Mamba-2, parallel heads
channels 2 : 1 : 5 (SSM : attention : MLP)
34B · 18T tokens
256K context; more attention channels ‘significantly degrades’ quality
Hunyuan-TurboS Tencent · 21 May 2025
Mamba-2
AMF/MF blocks, 128 layers
560B · 56B active
16T tokens, 256K context; described as the first industry-deployed large-scale Mamba model
Qwen3-Next Alibaba · 11 Sep 2025
Gated DeltaNet
3 : 1 (36 : 12 of 48)
80B · 3B active
15T tokens at under 10% of Qwen3-32B’s training cost; over 10× its throughput beyond 32K
Granite 4.0 IBM · 2 Oct 2025
Mamba-2
9 : 1
32B · 9B active
Over 70% less RAM at long context; attention kept for in-context learning
Kimi Linear Moonshot · 30 Oct 2025
Kimi Delta Attention
3 : 1 (ablated 1:1, 3:1, 7:1, 15:1)
48B · 3B active
5.7T tokens; KV cache cut by up to 75%; up to 6× decoding throughput at 1M context
Nemotron 3 Nano NVIDIA · 15 Dec 2025
Mamba-2
6 attention of 52 (23 Mamba-2, 23 MoE)
30B · 3B active
25T tokens, 1M context; up to 3.3× the throughput of GPT-OSS-20B and Qwen3-30B-A3B
Qwen3.5-397B-A17B Alibaba · 16 Feb 2026
Gated DeltaNet
3 : 1 (45 : 15 of 60)
397B · 17B active
262K native context, 1M extended; Apache 2.0
Nemotron 3 Ultra NVIDIA · 4 Jun 2026
Mamba-2
select attention layers (count not on the card)
550B · 55B active
1M context; latent MoE; multi-token prediction for speculative decoding
Kimi K3 Moonshot · 27 Jul 2026
Kimi Delta Attention
3 : 1 (69 : 24 of 93)
2.8T · 104B active
1M context; first open-weight model at the frontier by its publisher’s benchmarks
Nemotron 3.5 Lightning NVIDIA · 11 Aug 2026
Mamba-2
6 attention of 52 (23 Mamba-2, 23 MoE)
31.6B · 3.6B active
1M context; OpenMDW licence; runs on one GPU
Qwen3.8-Max Alibaba · 12 Aug 2026
Gated DeltaNet
3 : 1 (69 : 23 of 92)
2.4T · 95B active
262K native context, 1M extended; largest open hybrid
Table 1. Hybrid models with state space or linear-attention layers, March 2024 to August 2026, from the papers and model cards cited in the text. No large model in the period shipped with no attention; none kept more than a quarter of its layers as attention. The linear-attention family settled on 3:1, the Mamba-2 family nearer 1 in 10.
The ratio question
Why roughly one attention layer in four? The honest answer is that it is one in four in the linear-attention family and nearer one in ten in the Mamba-2 family, and the evidence for either is empirical. Kimi Linear’s report (30 October 2025, arXiv 2510.26692) is the most explicit: at 1:1, 3:1, 7:1 and 15:1, training loss was nearly flat and validation perplexity was best at 3:1, 5.65 against 5.70 at 7:1 and 5.82 at 15:1, while 1:1 validated the same as 3:1 and cost more to serve.
Qwen3-Next, Qwen3.5, Qwen3.8 and Kimi K3 all chose 3:1. Nemotron-H and Nemotron 3 chose about 8 percent, Granite 4.0 chose 9:1, MiniMax-01 chose 7:1, and Falcon-H1, which runs attention and Mamba-2 heads side by side in the same layer, gave attention the smallest share of channels and found that giving it more hurt. The ratio is a budget: enough global lookups to copy, retrieve and follow a format, and no more, because every attention layer brings its cache back with it.
What the attention layers are kept for is now stated in the papers rather than guessed. MiniMax found that pure lightning attention could not retrieve. IBM kept transformer blocks for in-context learning, few-shot prompting in particular. Kimi Linear removes positional encoding from its attention layers entirely and lets the KDA layers carry position, which tells you the division of labour its designers see: the linear layers compress and order the stream, the attention layers look things up in it.
Kimi K3 goes one step further and adds an extra attention layer at the very end, so that the last thing the model does before predicting a token is a global lookup. Of the three failure modes of a fixed state, two are retrieval in some form. One attention layer in four is the price of retrieval.
The ecosystem argument, dated
If the architecture argument was over by mid-2024, the ecosystem argument explains the next two years. Hugging Face’s transformers library added Mamba on 21 March 2024. llama.cpp merged Mamba on 8 March 2024, and the pull request to support Jamba, opened on 25 May 2024, merged on 9 July 2025, fourteen months later, because a hybrid needs a recurrent state stored separately from the KV cache and nothing in the engine assumed that. vLLM added Jamba in version 0.5.1 on 5 July 2024 as its first state space model, pure Mamba in 0.6.3 on 14 October 2024, and Falcon Mamba in 0.6.4 a month later; hybrids remained, in the vLLM team’s own description, a fragile arrangement in which the Mamba state was allocated outside the paged cache and the operator had to guess the maximum number of sequences, until a rebuild described on 5 November 2025 gave them a unified allocator, CUDA graphs, and prefix caching marked experimental.
TensorRT-LLM added Mamba-hybrid support in 0.19.0 on 9 May 2025 and Nemotron-H in 0.20.0 on 19 June. SGLang’s 0.5.2 on 12 September 2025 shipped Qwen3-Next support, a Mamba kernel and a flash-linear-attention kernel in the same release, the day after the model. Every date in that paragraph marks a quarter in which an enterprise could not serve a hybrid on its standard stack, and the dates run to within a year of today.
The costs that do not show up in release notes are larger. When NVIDIA benchmarked Nemotron-H in March 2025 it ran its own model on an initial Megatron implementation and the transformer baselines on vLLM 0.7.3, and said so, because there was no mature engine for its own architecture. IBM shipped a dense transformer version of Granite 4.0 alongside the hybrids for platforms that did not yet support them.
Quantisation recipes, fine-tuning libraries, speculative decoding, disaggregated prefill: every optimisation a platform team relies on was built for a cache and had to be rebuilt for a state. And the people. A decade of practitioners can read an attention map, debug a KV cache and tune a FlashAttention kernel, and most have never looked at a scan.
MiniMax’s account of why M2 went back to full attention, published with the SGLang team on 4 November 2025, is the ecosystem argument from the inside. Hybrids matched full attention on standard benchmarks and showed deficits in multi-hop reasoning at scale; a sliding-window hybrid fell from 90 to 72 on a 128,000-token retrieval test; linear kernels were memory-bound even in training; low-precision state storage, prefix caching and speculative decoding were open problems.
The conclusion was not that efficient attention is wrong but that its benefits ‘will eventually emerge’ once the evaluations and the infrastructure catch up. Seven months later, on 27 May 2026, MiniMax announced that M3 would use block-sparse attention, with a reported 15.6 times faster decoding at a million tokens than M2. Not linear. Sparse. That is a third answer to the same problem, and it is the one Google chose for Gemma 4 in April: softmax attention in every layer, five of six layers restricted to a 1,024-token window, no linear layers at all.
The Mamba paper itself was rejected by ICLR in January 2024, on reported scores of 8, 8, 6 and 3, with one reviewer asking for a comparison against a 10-billion-parameter transformer. The architecture argument was already won by then. The ecosystem had not read the paper.
The 2026 evidence
Here is what I can verify as of today, 8 October 2026. Qwen3.5-397B-A17B, released on 16 February under Apache 2.0, has 60 layers in fifteen groups of three Gated DeltaNet layers and one attention layer. NVIDIA’s Nemotron 3 family, announced on 15 December 2025, is hybrid Mamba-2, attention and mixture of experts throughout: Nano with six attention layers in 52, Super at 120 billion parameters on 11 March 2026, Ultra at 550 billion with 55 billion active on 4 June, all with a million-token context.
Mamba-3 appeared at ICLR 2026 and reports that at 1.5 billion parameters it beats Gated DeltaNet by 0.6 points, 1.8 with its MIMO variant, and matches Mamba-2’s perplexity with half the state. Kimi K3, 27 July: 69 KDA layers and 24 Gated MLA layers. Nemotron 3.5 Lightning, 11 August: 23 Mamba-2 layers, 23 expert layers, six attention layers. Qwen3.8-Max, 12 August: 69 Gated DeltaNet layers and 23 attention layers in 92.
Qwen3.8-Flash-Next, described in a paper of 31 August (arXiv 2608.30320), keeps one attention layer in four and swaps those for a sparse attention during continued pretraining, and Alibaba says the design previews Qwen4.
What I cannot verify matters as much. Google has not disclosed the architecture of Gemini 4 Argon, announced on 30 September. Gemma 4’s technical report, posted on 2 July, uses sliding-window and global softmax attention and nothing linear. I found no Meta disclosure of linear or state space layers in any 2026 Llama.
So the picture is this: every Chinese frontier lab that publishes weights has a linear layer in three positions out of four; NVIDIA and IBM have Mamba-2 in most positions; the American closed labs say nothing; Google’s open models and MiniMax took the sparse route. The pure state space model is nowhere on that list. Its layer is in most of it.
What it means for an enterprise platform
Which workloads benefit. Anything where the context is long and the lookups are few: a 400-page contract summarised, a quarter’s call transcripts classified, an agent trace that has grown to 300,000 tokens of tool output by step forty, speech, where a per-token cost that does not grow is the only kind that survives an hour of audio. In every one of those the KV cache is the bill, and in a 3:1 hybrid it is a quarter of the size.
Anything where the task is exact copying from far back, a phone number from page 90, a stack trace reproduced verbatim, a figure that must match the source to the digit, is where the fixed state is weakest, and the evaluation bench has to include a probe for it before the router sends work there. Hybrids are not worse models. They are models with a known weak spot, and the weak spot is the one retrieval-augmented pipelines exercise hardest.
What to ask a vendor. Three questions, and the answers belong in the model registry beside the router configuration. First, what fraction of layers are attention, and what is the state size of the others; a vendor who will not say is selling you a cache you cannot size. Second, how is long context priced.
Anthropic’s current rate card gives Claude 4.6 and later models a million-token window at the standard rate, but prices Haiku 5.5 at five times the rate for prompts over 100,000 tokens, $0.50 against $0.10 per million input and $2.50 against $0.50 output; Google’s page, updated yesterday, still doubles the input price of Gemini 2.5 Pro and 3.1 Pro above 200,000 tokens. Those tiers are an architecture decision priced in public: a provider whose attention grows with context charges for it, and the September price war did not touch them.
Third, which engine serves the model, at which version, with prefix caching verified on the state, because a hybrid without a working prefix cache throws away the single largest saving in agent workloads, as the July post on tokens to done argued.
A worked example: one million tokens, three ways
Take a reference design, mine rather than any vendor’s, with every assumption labelled. Sixty-four layers. Attention layers use grouped-query attention with 8 key-value heads of dimension 128, the Llama 3 arrangement, in 16-bit precision. Each attention layer stores one key and one value per token: 2 × 8 × 128 × 2 bytes, 4 kilobytes per token per layer. Linear layers keep a state of 32 heads, each a 128-by-128 matrix, in 16-bit: 32 × 128 × 128 × 2 bytes, 1 megabyte per layer, whatever the length, plus a short convolution state small enough to ignore. Those head counts are the ones Qwen3-Next and Kimi Linear publish.
Design A, all 64 layers attention: 256 kilobytes of cache per token. One sequence of 128,000 tokens is 32 gigabytes; 512,000 is 128; a million is 256 gigabytes, which is more memory than three H100s hold, for one user, before a single weight is loaded. Design B, all 64 layers linear: 64 megabytes per sequence at any length, so a million-token context costs what a ten-token one costs, and a thousand concurrent million-token users fit in 64 gigabytes.
Design C, the 3:1 hybrid, 16 attention and 48 linear layers: 64 kilobytes per token plus 48 megabytes of state, so 8 gigabytes at 128,000 tokens, 32 at 512,000, 64 at a million. Figure 2 draws it.
The arithmetic says three things. The hybrid is a quarter of the pure transformer at every length, which is exactly the 75 percent cache reduction Kimi Linear reports, and because throughput at long context is bounded by how many sequences fit in memory, a quarter of the memory is roughly four times the batch, which is most of the six-fold decoding speedup at a million tokens in the same paper.
The pure linear design is a thousand times smaller again, and that thousand-fold has bought nothing anyone was willing to ship, because the 16 attention layers in Design C are where the copying and the recall live. And at 128,000 tokens, the context most enterprise work actually uses, the hybrid’s 8 gigabytes per sequence is the difference between a GPU serving ten users and forty. vLLM’s engineers measured the same thing on a real model, Nemotron Nano 2, and put the attention cache at nearly 200 times the Mamba state at 128,000 tokens.
What a provider charges for long context is, to a first approximation, which of these three lines it is on. The Kimi K3 post priced the weights; this is the memory the weights need to be useful.
128K tokens
GiB of cache or state per sequence, 16-bit
All attention
32
3:1 hybrid
8.05
All linear
0.06
512K tokens
GiB of cache or state per sequence, 16-bit
All attention
128
3:1 hybrid
32.05
All linear
0.06
1M tokens
GiB of cache or state per sequence, 16-bit
All attention
256
3:1 hybrid
64.05
All linear
0.06
Figure 2. Serving memory for one sequence under the three designs in the worked example. Assumptions, all mine: 64 layers; attention layers with 8 key-value heads of dimension 128 (4 KiB per token per layer); linear layers with a 32 × 128 × 128 state (1 MiB per layer); 16-bit everywhere; 1M means 1,048,576 tokens. Bars are scaled within each panel; the all-linear bar is drawn at a minimum width so that it is visible at all. The hybrid tracks Kimi Linear’s reported 75 percent cache reduction at a 3:1 ratio.
What to do
Put architecture on the bench. Add three probes to the evaluation suite: a phone-book lookup at 50,000 and 200,000 tokens, a verbatim copy of a 2,000-token passage from the start of a 128,000-token prompt, and a format-following task with an unusual answer schema. Hybrids fail these first, and benchmark averages hide it.
Route by context length and retrieval load, not by vendor. Long documents, transcripts and agent traces go to the hybrid; exact copying and multi-hop retrieval go to full or sparse attention. The router already has the task class; add the context-length distribution.
Budget the cache, not the parameters. For every model in the registry record kilobytes of cache per token and megabytes of state per sequence, and size GPUs from those two numbers and the context lengths of real traffic.
Ask the three vendor questions and write down the answers. Attention fraction and state size; long-context pricing tiers; the engine version with prefix caching verified on the state.
Do not self-host a hybrid on an engine younger than its own prefix-cache support. The dates above are the warning. The rebuild that made hybrids first-class in vLLM is eleven months old, and the first engines to serve a new hybrid are routinely a week behind the weights.
Treat the ratio as a disclosure you can audit. Three to one, nine to one, or all attention: it is the number that tells you, before any benchmark, what a model will be good at, what it will be cheap at, and where to look for the failure.
Mamba did not replace the transformer, and the people who wrote it never quite said it would; the hybrid was in H3 a year earlier. What happened instead is that the industry ran the experiment at every scale it could afford, found the same quarter of attention it could not remove, and rebuilt the frontier around that ratio while the pure architecture lost every argument about tooling. The pure state space model lost the ecosystem argument. The state space layer won the architecture argument, by stealth, from inside the models it was meant to replace.
If you are buying inference in October 2026, three-quarters of what you are paying for is already a recurrence. The question worth asking a vendor is what the other quarter costs.
Ashish KumarHead of Platforms, AI & Data at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn