Built for the Real World · Essay · Unit economics

The price of intelligence halved in a week. Your architecture should notice.

Anthropic cut Opus on Monday. OpenAI answered within hours with GPT-6 Sol and Luna at half their predecessors’ prices. Sonnet 5.5 arrived on Sunday at a price that has not moved in a year. A price war is good news only if your system can take advantage of it.

Six descending steps, the tallest filled in clay and the rest outlined in slate, representing falling model prices
Each step is a price cut. Illustration by the author.

Here is the week in list prices, per million tokens, input and output. On Monday 22 September, Anthropic released Claude Opus 5.5 at $4 and $20, down from Opus 5’s $5 and $25, with cache reads cut from $0.50 to $0.20 and batch rates of $2 and $10. Anthropic’s headline was that it performs at Claude Fable 5.1’s level on most work at roughly 40 percent lower typical cost. Within hours, OpenAI released GPT-6 Sol at $2 and $10 and GPT-6 Luna at $0.10 and $0.50, each about half of what the GPT-5.6 equivalents cost on promotional pricing, with cached input at $0.20 and $0.01 respectively. On Sunday 28 September, Anthropic shipped Sonnet 5.5 at $2 and $10, the price Sonnet has carried for a year, and reported it within two points of Opus 5.5 on its knowledge-work and coding benchmarks while beating it on Terminal-Bench.

Three weeks earlier, GPT-6 Astra had launched at $10 and $50. On Monday morning the cheapest frontier-lab model was a hundred times cheaper than the most expensive one on input, and the gap between the flagship tier and the tier below it was a factor of two to five, for a capability gap that the labs themselves measure in low single digits. This is a price war, and like all price wars it rewards the customers who are set up to switch and punishes the ones who signed a contract for the model they liked in June.

What each release actually changed

Opus 5.5. The 20 percent cut on standard rates is the smaller half of the story. The larger half is the cache-read cut of 60 percent, from $0.50 to $0.20, because long-running agents re-read their context on every step and pay mostly for cache reads. The “40 percent lower typical cost” figure combines the price cuts with token efficiency: Opus 5.5 uses fewer tokens per task than Opus 5. But there is a detail in the fine print that any cost model has to respect. The 40 percent comparison uses Opus 5.5’s new default effort setting of medium against Opus 5’s default of high, while the benchmark table reports results mostly at max effort. The savings and the scores are measured at different settings. The model also ships five breaking changes, among them that thinking can no longer be disabled and that forced tool use is rejected, so it is not a drop-in for every workload. Context is one million tokens, output 128,000, with 300,000 available through the batch API.

GPT-6 Sol and Luna. Both carry a 1.05-million-token input window and 128,000 output. OpenAI attributes the halving to caching and inference improvements, and the benchmark story it tells is about cost per task rather than peak scores: on two benchmarks, DeepSWE and OSWorld, GPT-5.6 Sol’s maximum scores still exceed GPT-6 Sol’s, at much higher cost per task. Sol’s factual error rate is about half of GPT-5.6’s at comparable effort. The more interesting number is Luna’s. At maximum effort, Luna matches Sol’s DeepSWE score of 66.6 percent for roughly one-fifth of the cost, which is why OpenAI positions it for high-volume agents. Neither model touches Astra, which keeps about 73.5 percent on computer use against Sol’s 64.4.

Sonnet 5.5. Same price, different model. On Anthropic’s GDPval-AA, the professional knowledge-work benchmark, Sonnet 5.5 scores 1,844 against Opus 5.5’s 1,846. On CursorBench, agentic coding, 55.5 percent against 57.8. On Terminal-Bench 4.0, 70.6 percent against Opus 5.5’s 66.4 at its highest effort. Output generation is over 30 percent faster than Sonnet 5, and because it uses fewer tokens for the same work, Anthropic reports cost per task down by up to 30 percent with no change to the rate card. A model that says less to achieve the same result is a price cut you will not find on an invoice.

Input price

US dollars per million tokens, list, 29 Sep 2026

GPT-6 Astra
$10
Fable 5.1
$10
Opus 5.5
$4
GPT-6 Sol
$2
Sonnet 5.5
$2
GPT-6 Luna
$0.10

Cached-input price

the rate that dominates agent bills

Fable 5.1
$0.25
Opus 5.5
$0.20
GPT-6 Sol
$0.20
GPT-6 Luna
$0.01

Change this week

versus the model each one replaces

Opus 5.5 list
−20%
Opus 5.5 cache
−60%
GPT-6 Sol
−50%
GPT-6 Luna out
−58%
Sonnet 5.5
0% (−30% per task)
Figure 1. Clay: Anthropic. Slate: OpenAI. List prices only; promotional, batch and regional rates differ. The cached-input column is the one to watch for agents. Sonnet 5.5 held its price while closing most of the gap to Opus, which is a cut in all but name.

Why list price is the wrong number

Nobody pays list price per token; they pay per task. Four things in this week’s releases move cost per task more than the headline rates, and a budget that tracks only the headline will be wrong in both directions.

Cache reads. An agent that re-reads a 50,000-token context on every one of forty steps processes two million tokens of context, of which 1.95 million are cache hits. At $0.50 that context costs $0.98 a run; at $0.20 it costs $0.39. The 20 percent cut on fresh input saved a few cents. The cache cut saved sixty. Anthropic made the same move on 1 September when it cut Fable’s cache reads by 75 percent and estimated the saving at 25 percent for typical workloads and 45 percent for agentic ones. Cache pricing is now the lever the labs pull when they want to compete for agents specifically.

Effort defaults. Every current model exposes an effort dial, and every price comparison depends on where it is set. Opus 5.5’s default moved from high to medium. OpenAI’s own analysis of Sol says that effort beyond xhigh buys 0.8 to 2.2 points on coding tasks at 1.6 to 2.7 times the cost. A workload that was tuned for Opus 5 at high effort and migrated to Opus 5.5 at the new default is a different workload, cheaper and possibly worse, and nobody will notice unless the evaluation bench runs.

Token efficiency. Sonnet 5.5’s 30 percent cost-per-task reduction at unchanged prices, Gemini 3.6 Flash’s 17 percent fewer output tokens in July, Opus 5.5’s fewer tokens per task: the labs are now competing on how little the model says. This shows up nowhere on a rate card and everywhere on a bill.

Tier compression. When the second tier scores within two points of the flagship, the question is no longer which model is best but which tasks actually need the top two points. For most enterprise workloads the honest answer is very few, and the few that do are the ones a human will audit anyway.

A worked example

Take a workflow we run in several subsidiaries: triage and first-draft response for inbound supplier correspondence. Call it 200,000 cases a month. Each case has about 6,000 tokens of context, a 1,500-token retrieved policy block that is identical across cases and therefore cacheable, and produces about 400 tokens of output. Three routing policies, priced at this week’s list rates:

Routing policyModel mixMonthly tokensMonthly cost at listNotes
Flagship everywhere
the June default
100% Opus 5.51.2B input (0.3B cached), 80M output≈ $5,300$3,600 fresh input, $60 cached, $1,600 output. Same job on Opus 5 in June: ≈ $6,800.
Tiered by task class
router with three classes
70% Luna for classification and routing; 25% Sonnet 5.5 for drafts; 5% Opus 5.5 for escalationssame≈ $1,050Luna handles the bounded decisions at $0.10; Sonnet drafts at $2; Opus only where a human will review.
Tiered plus cached policy block
as above, with prompt caching
same mix, policy block cachedcached share rises to 25%≈ $880A further 16% from a one-line change. At June’s cache prices this step was worth half as much.
Table 1. Illustrative monthly cost for a 200,000-case correspondence workflow at 29 September list prices. The arithmetic is rough and the token counts are ours; the point is the ratio. Routing by task class is worth five times more than any single price cut, and the cache cut doubled the value of a change we had already made.

The flagship-everywhere policy got 22 percent cheaper this week without anyone touching it. The tiered policy got about 35 percent cheaper, because Luna’s and Sonnet’s cuts compound. The difference between the two policies is $4,000 a month on one workflow in one subsidiary, and we have several hundred such workflows. Price cuts are worth far more to an architecture that can route than to one that cannot, which is the whole argument of this post.

The supply side, briefly

It is worth asking why prices are falling this fast, because the answer tells you whether it continues. Part of it is competition between the labs, and this week’s sequencing, Anthropic in the morning and OpenAI by evening, makes that explicit. Part of it is engineering: OpenAI credits caching and inference improvements; both labs now ship models that produce fewer tokens per task. And part of it is capacity arriving. Alibaba this week unveiled its Zhenwu V900 chip, said Qwen 4 is being trained toward five to ten trillion parameters and set a target of twenty gigawatts of data-centre capacity by 2032; Anthropic has committed on the order of $80 billion to compute; NaiveAI released a 309-billion-parameter mixture-of-experts model with 15.5 billion active parameters and a native one-million-token context at a low price. More supply, more sparsity and more competition all point the same way. A tracker computed the spread across the top fifteen models this month at 119 times, with a new floor of $0.10 per million for a top-five model. I would not plan on the floor rising.

What I would plan on is volatility. Three labs now publish prices with expiry dates: Gemini 3.8 Flash doubles on 1 January, GPT-5.6 Sol’s promotional rate ends on 21 November, and Anthropic cancelled a scheduled Sonnet 5 increase on 11 August. A rate card is now a time series.

What an architecture that can take advantage looks like

  1. Every call goes through a router with a task class. Classification, extraction, routing and scoring go to the cheapest tier that passes the evaluation bench; this is also where the new single-pass decision models belong. Drafting and synthesis go to the middle tier. Planning, code and anything a human will audit go to the flagship. The class is set by the task; the model is set by configuration; a price change is a configuration change.
  2. Cost per task is a first-class metric. Log tokens by type, fresh, cached, batch and output, multiply by the rate card in force on the day, and attribute the result to the task and the workflow. If you cannot answer “what did this workflow cost yesterday versus today” in one query, you cannot act on a price cut, and you will not notice an effort-default change either.
  3. Effort is a configuration value, not a default you inherit. Pin the effort level per task class and re-evaluate it when the vendor changes its default. A silent move from high to medium is a quality change disguised as a saving.
  4. Re-run the bench on release day. This week moved the cost-performance frontier twice in six days. A bench of real tasks that runs in hours tells you whether Sonnet 5.5 can take a workload from Opus 5.5 before finance asks why the bill did not fall.
  5. Contract for flexibility, not volume. Compute commitments that lock spend to one vendor’s rate card are now a liability. OpenAI’s own DevDay marketplace lets enterprises apply commitments to open models via third parties, which is an admission that customers want portability. Negotiate for it.
  6. Budget with expiry dates. Every promotional rate enters the cost model with its end date. The January doubling of Gemini 3.8 Flash is a line item in next year’s plan, not a surprise in February.

The effort-default trap, in detail

The subtlest change this week is the one most likely to cause an incident, so it deserves its own section. Every current frontier model exposes an effort control, five levels on Anthropic’s models from low to max, a similar ladder on OpenAI’s, and the labs are now quietly moving the defaults. Opus 5.5 ships with a default of medium where Opus 5 defaulted to high. Anthropic’s 40 percent cost claim is computed across that change. Its benchmark table is computed mostly at max. Both numbers are true; neither describes a workload you are likely to run unchanged.

Consider what happens to a workflow that migrates by changing one string in a configuration file, which is what we tell people to do. It was tuned on Opus 5 at high effort. It now runs on Opus 5.5 at medium. The bill falls, satisfyingly. The accuracy on hard cases falls too, by an amount nobody measured because the migration was a one-line change and the bench was not re-run. The first sign is a downstream complaint, weeks later, that the agent has started missing a class of exceptions it used to catch. By then the cheaper bill has been reported as a saving and the regression has been attributed to something else.

OpenAI’s own guidance on Sol makes the trade-off explicit in the other direction: effort above xhigh buys 0.8 to 2.2 points on coding tasks for 1.6 to 2.7 times the cost. That is a curve, and a workload should sit at a chosen point on it, not at whatever point the vendor set this month. Our rule is now that effort is pinned per task class in the router’s configuration, never inherited from the model’s default, and that any vendor change to a default triggers a bench run before the migration ticket closes. This is boring. It is also the difference between a 35 percent saving and a 35 percent saving with an unexplained quality drop attached.

Cache design is now an engineering discipline

If cached input is the rate that dominates agent bills, then prompt structure is a cost control, and it has rules. A cache hit requires that the prefix of the prompt be byte-identical to a prefix the provider has seen recently. Anything that varies early in the prompt, a timestamp, a session identifier, a user name, a randomly ordered list of tools, breaks the hit for everything after it. The design that works is stable-prefix, variable-suffix: system instructions first, then reference material that does not change across calls, then tools in a fixed order, then the conversation, then the current task last. Agents that compact or rewrite their history on each step defeat the cache by construction; agents that append do not.

The economics are large enough to justify re-architecting. Luna’s cached-input rate is a tenth of its fresh rate; Opus 5.5’s is a twentieth. An agent whose prompt is cache-friendly pays the fresh rate on perhaps five percent of its input tokens. One whose prompt is not pays it on everything. For the same workload, the same model and the same rate card, that is a ten- to twenty-fold difference in the input line, and the input line is most of the bill. We now review prompt structure for cacheability the way we review database queries for indexes, and for the same reason.

The same discipline applies to batch. Both labs now price batch processing at half the standard rate, and OpenAI’s batch rates for Luna fall to $0.05 and $0.25. Any workload that does not need an answer in seconds, which is most back-office automation, belongs on batch. The organisational obstacle is that engineers build for interactive latency by default. The fix is a router that knows which task classes are latency-tolerant and sends them to the batch endpoint without the application having to decide.

The objection: cheaper tokens, bigger bills

The counterargument is Jevons: when a resource gets cheaper, consumption rises faster than the price falls, and total spend goes up. In AI this is already visible. Agents that were too expensive to run continuously become viable at Luna’s prices, and a continuously running agent consumes more tokens in a week than a chat assistant does in a year. OpenAI’s Dots, which run always-on, are a bet on exactly that. Cheaper intelligence will almost certainly mean larger AI budgets, not smaller ones.

That is not a reason to ignore the price war. It is the reason to take it seriously. If total spend is going to rise because new workloads become viable, the unit economics of each workload are the only thing between a budget that grows with value and one that grows with waste. The organisations that come out of this well are not the ones that picked the right model in September. They are the ones whose systems can move when the price does, and whose cost accounting can tell the difference between a cheaper token and a more expensive habit.

Sources

  1. Anthropic. Claude Opus, pricing as published 22 September 2026.
  2. Digital Applied. Claude Opus 5.5: pricing, benchmarks and breaking changes. 22 September 2026.
  3. Digital Applied. GPT-6 Sol and Luna: API prices, benchmarks and trade-offs. 22 September 2026.
  4. MarkTechPost. Anthropic releases Claude Sonnet 5.5 at the same $2/$10 price. 28 September 2026.
  5. Local AI Zone. September 2026 AI model updates.
  6. The Neuron. Everything that happened in AI today, 28 September 2026: Alibaba’s Zhenwu V900 and capacity targets; NaiveAI’s Naive-N0.5-Flash.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn