On Tuesday 21 July, Google DeepMind released three models under the Gemini brand, and the sentence it chose to lead with was not a benchmark. It was that Gemini 3.6 Flash “reduces output token usage by 17 percent compared to 3.5 Flash.” Two weeks earlier, on 9 July, OpenAI had made GPT-5.6 generally available as three tiers, Sol, Terra and Luna, after a review by the US Department of Commerce, the first model launch to go through one. On 15 July, Thinking Machines released its first model, Inkling, 975 billion parameters under the Apache licence. On 17 July, Moonshot announced Kimi K3 at 2.8 trillion parameters, and published a price list on which the same input token costs $0.30 if it is a cache hit and $3.00 if it is not.
Taken together, July is the month the labs stopped competing only on how smart a model is and started competing on how much it costs to get a task done. That is a different contest with different winners, and it changes what an enterprise should measure. I run an AI platform whose bill is dominated by agents, loops that call a model dozens of times per task, and the number I care about has never been intelligence per token. It is tokens to done.
What shipped, and what each release is really selling
Gemini 3.6 Flash is Google’s new default Flash-tier model at $1.50 per million input tokens and $7.50 output, with the same 1,048,576-token input window and 65,536-token output limit as 3.5 Flash. Google reports a 65 percent improvement on DeepSWE, the autonomous software-engineering benchmark, and code-editing precision of 49 percent against 37. But the claim it chose to emphasise is that the model needs “fewer reasoning steps and tool calls” for multi-step work and emits 17 percent fewer output tokens for the same result. For a chat assistant that is a rounding error. For an agent that runs a forty-step loop, output tokens are the expensive line, and 17 percent fewer of them at every step compounds into a materially cheaper task.
Gemini 3.5 Flash-Lite is the throughput tier: $0.30 per million input and $2.50 output, 350 output tokens per second, and, by Google’s account, better than 3 Flash on agentic benchmarks including SWE-Bench Pro at 54.2 percent against 49.6. The pitch is explicit: high-throughput, low-latency agentic workflows. This is the model for the bounded steps inside an agent, the classifications and routings and extractions that do not need the flagship and happen a thousand times a minute.
Gemini 3.5 Flash Cyber is a variant tuned to find and fix vulnerabilities, deployed only inside Google’s CodeMender agent for a limited pilot of governments and trusted partners. It is the first of what would become, by September, a pattern of gated security tiers across all the major labs, and I will return to it then.
GPT-5.6 arrived as a lineup rather than a model: Sol as the flagship at $5 per million input tokens, Terra at roughly half Sol’s price with what OpenAI describes as GPT-5.5-level intelligence, and Luna as the small, fast tier at $0.20 input and $1.20 output. The three-tier shape is itself the message. OpenAI is telling customers that most work does not need Sol, and that the right architecture routes between tiers. The Commerce Department review is a footnote to the pricing story but a headline in its own right: the series sat in limited preview from 26 June for government-approved partners before the broad launch, and that is a precedent that will not be reversed.
Kimi K3 is 2.8 trillion parameters with 16 of 896 experts active per token, a one-million-token context, and open weights promised for 27 July. The price list is the interesting part: $0.30 per million for cached input, $3.00 for uncached, $15.00 for output. A ten-fold spread between a hit and a miss on the same token is the clearest statement any lab has made that caching is now a first-class pricing dimension. Inkling, from Thinking Machines, is 975 billion parameters under Apache 2.0 and the first serious open-weight release from a US frontier-adjacent lab this year; I will treat the open-weights wave on its own when there is more of it to assess.
The anatomy of an agent’s bill
A chat request costs what it costs: some input, some output, done. An agent is different. It reads a task, calls a tool, reads the result, reasons, calls another tool, and repeats until it decides it is finished, and the model is invoked once per step with the whole history so far as input. Three quantities determine what that costs, and the July releases each move one of them.
Steps to done. How many times the loop runs before the agent declares success. Google’s claim that 3.6 Flash needs fewer reasoning steps and tool calls is a claim about this number. It is the most powerful lever, because it multiplies everything else: a loop that finishes in 25 steps instead of 40 cuts input, output and latency together.
Tokens per step. How much the model says each time. The 17 percent output reduction is this lever. Output tokens are typically five times the price of input, so a model that is terse where terseness is correct is a model that is cheaper at identical capability.
Price per token, split by cache. Each step re-sends the history, and most of the history is identical to the previous step, so most input is cacheable. Kimi K3’s ten-fold hit-versus-miss spread and OpenAI’s and Anthropic’s cached-input rates mean that the effective input price of a well-designed agent is a small fraction of the list price, and the effective price of a badly designed one, whose prompt changes in a way that defeats the cache, is the full rate.
Run the arithmetic for a real workload. One of our subsidiaries runs an agent that reconciles supplier invoices against purchase orders: about 60,000 tasks a month, forty steps each when it started, 6,000 tokens of context, 300 tokens of output per step. The naïve version of that loop on a Flash-class model costs about $27,000 a month. The version with caching, a model that finishes in fewer steps, and the bounded steps routed to a Lite-class tier costs about $2,500. Nothing about the task changed. The model got terser, the loop got shorter, and the architecture stopped paying flagship prices for classification.
Why “tokens to done” is the benchmark that matters
Public leaderboards score a model on whether it gets the answer. They do not, as a rule, score how many tokens it spent getting there, and the two are increasingly decoupled. A model that reasons for 4,000 tokens and a model that reasons for 800 can land on the same answer; the first costs five times more and takes five times longer, and the leaderboard calls them equal. Google’s decision to lead with a token-efficiency number is an acknowledgement that the labs’ customers have noticed.
The right internal benchmark is therefore the pair: accuracy on your tasks, and tokens consumed per task on your tasks, measured on the same runs. We have started reporting both, and the results reorder the leaderboard. On our invoice workload, a model that scores two points lower on accuracy and uses 40 percent fewer tokens is the better model, because the two points are recovered by the escalation path and the 40 percent is recovered by nobody. On our contract-analysis workload the opposite holds, because errors are expensive and the volume is small. The point is not that efficiency always wins. It is that you cannot make the trade-off without the second number, and almost nobody publishes it.
Five design rules that fell out of July
- Measure tokens to done, per task class, on your own workloads. Steps, input tokens split by cached and fresh, output tokens, cost at the rate card of the day. If you measure only accuracy you will pick the wrong model for most of your volume.
- Design prompts for the cache. Stable prefix, variable suffix; system instructions and reference material first, the changing task last; no timestamps or random identifiers in the cached region. A prompt that defeats the cache at Kimi K3’s prices costs ten times what it should, and the cheaper labs’ ratios are not far behind.
- Route bounded steps to the throughput tier. Inside most agent loops, more than half the steps are classification, extraction or a yes/no check. Flash-Lite, Luna and their equivalents exist for exactly those steps. Reserve the flagship for the steps that need judgment.
- Penalise verbosity in evaluation, not just in prompts. A model told to be concise will be concise until the task gets hard. A model that is terse by training stays terse. Prefer the second, and score it.
- Cap steps and make the agent justify extensions. An agent that can take unlimited steps will, on hard cases, take them. A step budget with an explicit “request more” action turns an open-ended cost into a bounded one and produces a signal about which tasks are genuinely hard.
Measuring tokens to done in practice
The metric is simple to describe and surprisingly awkward to implement, so here is how we did it and what went wrong.
The first attempt counted tokens from the provider’s usage fields on each call and summed them per task. It was wrong in two ways. It did not distinguish cached from fresh input, so a prompt change that defeated the cache showed up as a cost increase with no change in token count. And it attributed tokens to calls, not tasks, so an agent that retried a step three times looked like three cheap calls rather than one expensive step. The fix was a task identifier propagated through every call in the loop, a per-call record of input tokens split by cache status, output tokens, model, effort level and latency, and a daily rollup by task class. The rollup multiplies each line by the rate card in force on that day, which means the rate card is itself a versioned table, with effective dates, that someone maintains. In July, with three labs changing prices and two of them publishing promotional rates with end dates, maintaining that table is a real job.
The second attempt produced the right numbers and the wrong conclusions, because it compared models at their default effort settings. A model with a higher default looked more expensive and more accurate, and the comparison was really between two effort levels. The bench now pins effort per task class and compares models at the same pin. This is the same lesson the price war would teach again in September, and we learned it first in July from Gemini’s token-efficiency claim, which is only meaningful at matched effort.
The third thing that went wrong is organisational. Once tokens to done was visible per task class, teams started optimising it, and some optimised it by making agents give up earlier. A step cap that is too tight raises the escalation rate, and escalations are handled by people, whose cost is not in the token ledger. The metric has to be paired with a task-success rate and an escalation rate, reported together, or it will be gamed by a loop that finishes cheaply by finishing badly. The three numbers together, cost per task, success rate, escalation rate, are what we now report, and the decision rule is the cheapest configuration that holds success and escalation within the thresholds the business owner set.
The three-tier lineup as an architecture signal
OpenAI did not ship GPT-5.6 as one model with a price. It shipped Sol, Terra and Luna, with Terra positioned at half Sol’s price and Luna an order of magnitude below that. Google shipped Flash, Flash-Lite and a gated Cyber variant the same month. The lineup is the vendors’ recommended architecture, stated as a catalogue: most calls go to the cheap tier, some to the middle, few to the top. A customer that routes everything to the flagship is paying for capability the vendor itself assumes most calls do not need.
Inside an agent loop the tiers map to step types. The flagship plans, decomposes and handles the steps where judgment is required. The middle tier drafts, summarises and synthesises. The cheap tier classifies, extracts, routes, checks and scores, which, in every loop we have instrumented, is the majority of steps by count and a minority by difficulty. The router that implements this needs one thing the applications do not have: a declared task class on every call. Getting development teams to declare the class is the cultural work; once it is declared, the economics follow from configuration.
The tiers also make the evaluation bench tractable. Instead of asking which single model is best, the bench asks, per task class, which is the cheapest model that passes. That question has a different answer for classification than for synthesis, and it is the answer the lineup was designed to let you give.
The Commerce Department footnote
It would be wrong to let July pass without noting that the GPT-5.6 launch was the first to go through a government review before general release. The series sat in limited preview for government-approved partners for two weeks, OpenAI met with agencies, and only then did it ship broadly. The details of the review are not public. The precedent is. For a platform that serves regulated businesses, the practical implication is that the most capable models will increasingly arrive with a lag between announcement and availability, and that the lag will be filled by review processes whose criteria we do not see. Plan for the lag. Do not plan launches around a model that has been announced but not cleared.
July’s releases were not the most capable of the year. They were the ones that made the cost of capability legible. The labs now compete on tokens per task, cache economics and tiering, and they are publishing numbers for all three. An organisation that measures the same things can take advantage of every release. One that measures only accuracy will keep buying the flagship for work that a Lite model finishes faster, and will keep wondering why the bill does not fall when the prices do.