Built for the Real World · Essay · Model strategy

Opus 5 at the old price: an unchanged rate card is not an unchanged bill

Anthropic released Claude Opus 5 on 24 July at $5 and $25 per million tokens, the same price as Opus 4.8, with scores ahead of Fable 5 on several agentic evaluations. The rate card did not move. Tokens per task, the tokenizer and the model’s behaviour did, and the right way to receive a release is a bench run, not a migration.

Abstract illustration for “Opus 5 at the old price: an unchanged rate card is not an unchanged bill”

On Friday 24 July 2026, Anthropic released Claude Opus 5 at $5 per million input tokens and $25 per million output tokens, the price Claude Opus 4.8 has carried since 28 May. The context window is one million tokens, as both the default and the maximum, with 128,000 tokens of output. Anthropic’s own description is that the model “comes close to the frontier intelligence of Claude Fable 5 at half the price,” and its published results support a stronger claim: on several agentic evaluations Opus 5 is ahead of Fable 5. Within an hour the first migration tickets were open on our platform. By Saturday morning two teams had changed one string in a configuration file and marked the upgrade done.

This post is about why that is the wrong way to receive a model release, even a good one, which Opus 5 is. An unchanged list price is not an unchanged cost. The rate card is the one variable in a cost-per-task calculation that did not move on Friday. Three others did: how many tokens the model spends to finish a task at its default effort, whether the tokenizer counts the same text the same way, and how the model’s changed behaviour alters the retry rate. The only instrument that measures all three on your own work is an evaluation bench, which is why our rule is that a release is received with a bench run, not a migration.

What shipped

The specifications first. Opus 5 replaces Opus 4.8 as the Opus-tier model eight weeks after 4.8 shipped, under the model ID claude-opus-5. Prompt caching is $6.25 per million tokens for a five-minute cache write, $10 for a one-hour write and $0.50 for a cache read, all unchanged from 4.8; the minimum cacheable prompt drops from 1,024 tokens to 512. Batch requests are $2.50 and $12.50. Fast mode, still a research preview on the first-party API only, is $10 and $50, the same premium as on Opus 4.8; the same day’s release notes removed fast mode from Opus 4.7 altogether. The model is on the Claude API, Amazon Bedrock, Google Cloud and Microsoft Foundry from day one, and it is not subject to the 30-day data-retention requirement that Fable 5 carries.

Two API changes matter more than any of those numbers. Thinking is now on by default: on Opus 4.8 a request that omitted the thinking field ran without thinking, and on Opus 5 the same request thinks, with the effort parameter deciding how much. The default effort stays at high, on a ladder of low, medium, high, xhigh and max. And there is one breaking change: a request that sets thinking to disabled at effort xhigh or max now returns a 400 error. Anthropic’s prompting guide adds a third change, outside the API: prompts that add a final verification step should lose it, because Opus 5 verifies its own work unprompted and the instruction now buys over-verification, which is to say tokens.

Now the results, all as Anthropic reported them. On SWE-bench Verified Opus 5 scores 96.0 percent, against 95.0 for Fable 5 and 88.6 for Opus 4.8 at its May launch. On SWE-bench Pro, the harder set, it scores 79.2 against 80.0 for Fable 5 in Anthropic’s comparison and 69.2 for Opus 4.8. On OSWorld 2.0, the long-horizon computer-use benchmark, it reaches 70.57 percent against 55.7 for Opus 4.8, and Anthropic says that beats Fable 5’s best result at just over a third of the cost. On ARC-AGI-3 the ARC Prize Foundation verified 30.16 percent at high effort, where GPT-5.6 Sol scores 7.78 and Opus 4.8 scored 1.52. On Frontier-Bench v0.1, the 74-task successor to Terminal-Bench 2.1, Opus 5 scores 43.3 percent at max effort against 33.7 for Fable 5, 37.5 for GPT-5.6 Sol and 18.7 for Opus 4.8. On Zapier’s AutomationBench it passes 26.0 percent of tasks against 17.0 for Opus 4.8 and 17.4 for Fable 5.

SWE-bench Pro

percent resolved, Anthropic-reported, 24 Jul 2026

Opus 5
79.2
Fable 5
80.0
Opus 4.8
69.2

OSWorld 2.0

percent of computer-use tasks, same source

Opus 5
70.57
Fable 5
66.1
Opus 4.8
55.7

Frontier-Bench v0.1

mean reward, 74 tasks, max effort

Opus 5
43.3
Fable 5
33.7
Opus 4.8
18.7
Figure 1. Scores as reported at the 24 July release; clay is Opus 5. Fable 5’s SWE-bench Pro figure is the 80.0 in Anthropic’s Opus 5 comparison; its own launch post in June reported 80.3. Fable 5’s OSWorld 2.0 figure is as DataCamp read it from Anthropic’s charts. Frontier-Bench is an internal Anthropic run, mean reward over five attempts, with Opus 4.8 serving as the fallback on safety-classifier refusals. Every number here is a vendor number at a vendor-chosen effort.

Why “same price” hides three variables

A rate card prices a token. A budget pays for a task. Between the two sit three quantities a release can change without touching the rate card, and Opus 5 moves all three.

Tokens per task at the default effort. Anthropic’s guidance for Opus 4.8 was to start at xhigh for coding and agentic work. Its guidance for Opus 5 is to start at high, the default, to use low and medium “liberally” wherever evaluations show quality holds, and to run a fresh effort sweep rather than carry settings over. That is a different operating point, and the efficiency claims are stated at specific points on it: 26 percent fewer tokens than Opus 4.8 on legal work at max effort, a seventh of the reasoning tokens on a trading benchmark, a third fewer turns and tool calls on financial modelling. On AutomationBench, medium effort scores 24 percent at $0.89 a task against 26.0 at the headline setting. Good numbers, none of them yours. A route that ran Opus 4.8 at xhigh and now runs Opus 5 at the default high has changed two things at once, and a route that omitted the thinking field now pays for thinking it never bought before. The direction of the change cannot be read off the rate card. It is a measurement.

The tokenizer. This time the variable is zero, but it was not zero in April or in June. Opus 5 uses the tokenizer introduced with Opus 4.7 and shared with Opus 4.8 and Fable 5, so a prompt costs the same tokens on 5 as on 4.8. But Anthropic’s pricing page states that this tokenizer produces roughly 30 percent more tokens for the same text than the one before it, and Sonnet 5, released on 30 June, moved to it with the same warning. A route migrating from Sonnet 4.6 to Opus 5, or a fleet that mixes generations, pays about 30 percent more per word of input at an unchanged per-token price. The check is cheap: the token-counting endpoint over a fixed corpus of production prompts under the old model ID and the new, with the ratio written into the rate card next to the price. Ours, over 2,000 prompts, was 1.00 on Friday.

Behaviour, and the retry rate. This is the variable migrations miss, because it never appears in a specification; Anthropic’s prompting guide is candid about it. Opus 5’s user-facing responses run longer than earlier Opus models’, and lowering effort reduces thinking without reliably shortening the visible response. It narrates during agentic work. It delegates to subagents more readily, which pays off on large independent tracks and multiplies cost on small ones. It verifies and self-corrects unprompted, so legacy scaffolding that adds verification steps now compounds with the model’s own habit. Each of those is a token count the rate card cannot see, and each is tunable by prompt, which means the prompt written for 4.8 is now part of the cost of running 5.

Retries are the other half. A task that fails and is re-run costs the whole loop twice. Opus 5 should lower that rate on balance, since the pass rates are higher, but the system card, as the trade press reported it, notes slightly more factual hallucination than Opus 4.8, and in a coding loop a hallucinated library call is a failed test, which is a retry. Refusals are retries too. Opus 5’s cyber classifiers intervene about 85 percent less often than Fable 5’s, and in the Frontier-Bench run Opus 5 refused 5 percent of calls across 4 percent of trials where Fable 5 refused 42 percent of calls across 26 percent of trials, with Opus 4.8 as the fallback. A fallback is a retry on a different model, at that model’s rate, with the prompt cache invalidated by the switch. None of it appears on the invoice as a line called retries; all of it is in the cost per completed task.

What the gains mean in practice, and what they do not

For agentic coding the ten-point move on SWE-bench Pro, from 69.2 to 79.2, is the number that changes bills, because in a loop the pass rate divides the cost. I argued in Tokens to done on Thursday that an agent’s bill is three quantities multiplied together: steps to done, tokens per step and attempts to done. Opus 5 claims to move all three. If the claims hold on your tasks, the model is cheaper at the same price: a price cut that will not show on the rate card and will show, or not, on the bench.

Frontier-Bench deserves a closer look. The 43.3 percent is at max effort; Anthropic also reports 44.4 percent mean reward at xhigh, which is higher. More effort was not better on the longest-horizon tasks, consistent with what Anthropic’s effort documentation has said for two generations: beyond a point, max adds cost for small or negative gains. Pin effort per task class and sweep it, rather than running the flagship at the setting where the benchmark table was computed.

For computer use, the move from 55.7 to 70.57 on OSWorld 2.0 is the largest jump in the release and the one to treat most carefully. Opus 4.8 scored 83.4 percent on OSWorld-Verified at its launch in May and 55.7 on OSWorld 2.0 in July’s comparison: same model, two numbers, two benchmarks. A benchmark version is a variable like any other, and “computer use improved by fifteen points” without the version named is not a measurement. What the 2.0 number does say is that long-horizon desktop workflows have gone from failing more often than not to succeeding more often than not, which is the threshold at which a pilot becomes a product.

What none of the numbers mean is that your tasks will behave the same way. They are vendor-run, at vendor-chosen effort, on vendor harnesses, and the Frontier-Bench score includes work Opus 4.8 did on refused calls. I set out in The evaluation bench, in full why a public leaderboard does not transfer: the task distribution is someone else’s, the grader is unvalidated against your labels, and the cost was never recorded. The bench replaces those three with sampled production tasks, graders audited against humans and a dated rate card, and reports three numbers per task class. A model release is the event it was built for.

Opus 5 against Fable 5, after June

The comparison Anthropic invites is with Fable 5, and the one a platform must make includes what happened in June. Fable 5 launched on 9 June at $10 and $50, with cache reads at $1 per million, always-on thinking, and a 30-day data-retention requirement that rules it out under zero-retention contracts. On Friday 12 June, at 5:21 in the evening US Eastern time, Anthropic received a government directive under national security authorities and, in its words, had to “abruptly disable Fable 5 and Mythos 5” for all customers, because the order concerned foreign nationals and nationality could not be verified per request. Anthropic’s redeployment note attributes the order to export controls applied after Amazon researchers found a way past the model’s safeguards to surface software vulnerabilities. The controls were lifted on 30 June and access returned on 1 July, nineteen days after it went.

What followed matters for routing. Included access on subscription plans came back rationed, at up to half of weekly usage limits, with a deadline of 7 July after which use moved to prepaid usage credits; the deadline then moved to 12 July and to 19 July. Enterprise customers I spoke to in that fortnight were not debating capability. They were asking whether a model that a government can switch off for three weeks, and whose terms moved three times in a month, could be the only model behind a production route. From a platform’s point of view the answer is no, and it would be no for any supplier with that history. That is not a judgement on Anthropic, which was more transparent than most vendors would have been; it is the ordinary discipline of second-sourcing a critical input.

So the routing decision is not close. Opus 5 is the default for the Opus-class tier: half the price on every line, cache reads at half the rate, about 85 percent fewer classifier interventions, no retention constraint, and level with or ahead of Fable 5 on the agentic evaluations Anthropic publishes. Fable 5 stays as a routed exception for the task classes where a bench run shows a gap worth paying double for. The candidates are narrow: the longest-horizon autonomous work, where Anthropic itself still recommends Fable; the hardest SWE-bench Pro-like tasks, where 80.0 against 79.2 is inside any bench’s noise; and terminal-heavy work, where Fable 5 posted 88.0 on Terminal-Bench 2.1 at launch and Anthropic published no Opus 5 figure. The mechanics are in Routing by task class: the class is set by the task, the model by configuration. June adds one requirement: the switch from the exception back to the default has to be rehearsed, with the bench, before it is needed.

The release-day protocol, applied

Here is what happened on our platform between Friday afternoon and Tuesday. The first rule is that nobody changes a production route on release day; the migration tickets stay open and the configuration string stays what it was.

On Friday the rate card got a new entry dated 24 July, every rate above with its effective date and the note that fast mode is first-party only. The tokenizer multiplier was measured over the fixed corpus: 1.00 against Opus 4.8, 1.30 against Sonnet 4.6. Then the contract tests ran, every production request shape replayed against the new model ID, and two classes failed the way the release notes said they would: requests that disabled thinking at xhigh returned 400s, and requests that had omitted thinking now thought, so their max_tokens values, tuned for a model that did not think, were too tight. Both are fixes to make before any benchmark result means anything.

On Saturday any route inheriting the vendor default effort was pinned to what it had been running, so that the sweep compares like with like. The bench then ran on the held-out split, per task class, at the pinned level and one either side, plus a fourth run on the coding classes with the verification scaffolding deleted and a cap on subagent spawning. Graders ran pinned and the human audit sampled them as usual; a new model can change a judge’s error pattern as easily as a candidate’s. On Tuesday the report came back with three numbers per class, a person edited the routing table, and each migration ticket closed with the bench run identifier in it. Five working days for a release most teams absorb in five minutes, and at the end of it we know what the model costs us, not what it costs Anthropic.

CheckWhat to measureWhere the number comes fromPass condition
Rate card
day 0
List, cache write and read, batch and fast-mode prices, with the effective dateVendor pricing page and release notes of 24 JulyEntered in the dated rate card before any run is priced
Request contract
day 0
Every production request shape replayed against the new model IDIntegration suite; HTTP status and stop reason per requestZero 400s; max_tokens raised where thinking now runs
Tokenizer
day 0
Tokens for a fixed corpus, new model over oldToken-counting endpoint on the 2,000-prompt corpusRatio within 2 percent of 1.00, or a per-model multiplier goes in the rate card
Effort
day 1
The level in force on every routeRouter configuration, never the vendor defaultEvery route pinned; sweep runs at the pinned level and one either side
Quality
day 1 to 3
Score per task class on the held-out splitPinned graders, sampled by the human auditNo class regresses outside its noise band; audit kappa unchanged
Tokens to done
day 1 to 3
Fresh input, cache reads, cache writes and output per trajectoryUsage fields summed per task by the runnerCost per completed task at the pinned effort at or below the incumbent’s
Loop shape
day 1 to 3
Steps, tool calls, subagent spawns and response length per taskTrajectory logsWithin the class’s band, or the prompt is tuned and the run repeated
Refusals and retries
day 1 to 3
Refusal stop reasons, fallback switches, retry countStop reason and fallback blocks in responsesRefusal and retry rates at or below the incumbent’s, fallback cost included
Latency
day 1 to 3
p50 and p95 per classRunner timingsWithin the class SLO
Decision
day 4
The routing-table changeThe three-number reportEdited by a named person; bench run ID recorded in the migration ticket
Table 1. The release-day checklist as we run it. Day 0 is the vendor’s numbers entered and the contract verified; days 1 to 3 are the bench; day 4 is a person. Nothing on this list requires the vendor’s benchmark table to be right, and nothing on it is satisfied by the rate card being unchanged.

A worked example

Take one task class: a scoped bug fix in a service repository, where the agent is handed a failing test, finds the fault, patches it, runs the suite and repeats until it is green. It is one of the platform’s highest-volume coding-agent classes, about 40,000 tasks a month across the group. The token counts below are illustrative estimates, rounded, of the shape we see in this class; the prices are verified list rates. The point is the ratios, not the digits.

The incumbent: Opus 4.8 at xhigh, the setting Anthropic recommended for coding. A typical trajectory ran about 32 steps over a 40,000-token context that is almost entirely cache reads after the first step: roughly 50,000 fresh input tokens, 1.2 million cache-read tokens and 48,000 output tokens including thinking. At list that is $0.25, $0.60 and $1.20, about $2.05 per attempt. The bench pass rate on the class was 74 percent, so the cost per completed task was about $2.77.

The one-line migration: Opus 5 at the default high. Fewer turns, as promised, and less reasoning per turn: call it 24 steps, 50,000 fresh input tokens, 900,000 cache-read tokens and 36,000 output tokens, which is $0.25, $0.45 and $0.90, about $1.60 per attempt, at a pass rate of 80 percent: $2.00 per completed task, 28 percent below the incumbent. This is the result the migration ticket assumed, and the result that only appears if the prompt has been cleaned.

Like for like: Opus 5 at xhigh. Matching the old effort rather than the new default: 28 steps, the same fresh input, 1.05 million cache reads and 58,000 output tokens, about $2.23 per attempt at a pass rate of 84 percent, which is $2.65 per completed task: cheaper than the incumbent, better than it, dearer than the default. Whether four points of pass rate are worth 65 cents a task is a question for the task owner, not the router.

The one-line migration with the old prompt left in. The prompt written for 4.8 still says to add a final verification step and places no cap on subagents. Opus 5 obeys: 36 steps, 1.4 million cache reads, 64,000 output tokens, about $2.55 per attempt at the same 80 percent pass rate, which is $3.19 per completed task. That is 15 percent more than the incumbent, on an unchanged rate card, for a model that is better.

Four bills between $2.00 and $3.19 per completed task, from one price. The cheapest and the dearest are both the one-line migration; the difference is whether anyone read the prompting guide. Across 40,000 tasks a month the spread is about $48,000, on one class in one group, unknowable from the invoice until the month closes and knowable from the bench by Tuesday.

What to do

  1. Pin effort on every route. The vendor default is a value the vendor can change, and the guidance for it changed between 4.8 and 5. A route inherits nothing; it carries an explicit level, and a release triggers a sweep at that level and one either side.
  2. Give the rate card a date and a tokenizer column. Every model entry carries its list, cache, batch and fast rates with an effective date, and a measured token multiplier against a fixed corpus. Opus 5’s is 1.00 against 4.8 and about 1.30 against Sonnet 4.6; both belong next to the price.
  3. Receive releases with a bench run, and close tickets with its ID. No production route changes on release day. The migration ticket closes when it references a bench run on the held-out split, with three numbers per class and a named person’s edit to the routing table.
  4. Clean the prompt before you measure. On Opus 5, delete verification scaffolding, cap subagent spawning and state the response length you want. A bench run on the old prompt measures the old prompt’s cost on the new model, not the one you will run with.
  5. Route Opus 5 as the Opus-class default and Fable 5 as a routed exception. Half the price, fewer refusals, no retention constraint, and level or ahead on the published agentic evaluations. Keep Fable 5 for the classes where your bench, not the vendor’s, shows a gap worth double, and rehearse the switch back.
  6. Count refusals and fallbacks as retries. A refusal served by a fallback model is a second attempt at a second rate with a cold cache. A model that refuses 42 percent of calls on a hard benchmark is a different cost from one that refuses 5 percent, whatever the rate card says.
  7. Log tokens to done, by type, per trajectory. Fresh input, cache reads, cache writes and output, summed per task and attributed to the class. Without that log the three variables in this post are invisible, and a release that made a route cheaper and one that made it dearer look identical on the rate card, because they are.

Sources

  1. Anthropic. Introducing Claude Opus 5. 24 July 2026.
  2. Anthropic. Claude Platform release notes. 24 July 2026 entry.
  3. Anthropic. Pricing. Accessed 25 July 2026.
  4. Anthropic. Prompting Claude Opus 5. July 2026.
  5. Anthropic. Effort. July 2026.
  6. Anthropic. Redeploying Fable 5 and Mythos 5. 30 June 2026.
  7. MarkTechPost. Meet the New Claude Opus 5: Frontier-Class Agentic Coding and Computer Use at Unchanged Opus Pricing. 24 July 2026.
  8. DataCamp. Claude Opus 5 vs Claude Fable 5. July 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn