Routing by task class: a reference design for the enterprise model router
OpenAI put its updated GPT-5.5 Instant into the API today at the same $5 and $30 as GPT-5.5 itself, the sixth frontier rate card since April. Applications should not be choosing between them. This is the router design we use: a declared task class, a registry of constraints, cache-first prompts, half-price endpoints, fallback chains and a ledger that turns a price change into a config change.
This morning, 25 June 2026, OpenAI finished rolling out its updated GPT-5.5 Instant to every ChatGPT tier and made the same model available in the API under the chat-latest alias. Artificial Analysis lists it at $5 per million input tokens and $30 per million output, about 135 tokens per second, a time to first token of about 0.9 seconds and a 400,000-token context. Those are the rates OpenAI charges for GPT-5.5 itself, the reasoning model it shipped on 23 April. The same money buys a model that answers in under a second, or one that, at its highest effort setting, thinks for about 44 seconds before the first token arrives and scores 38 on the same intelligence index where Instant scores 26. Neither is the wrong model. The wrong thing is to ask an application developer to choose between them.
Ten weeks of releases sharpen the point. Claude Opus 4.7 on 16 April, with a new tokenizer. GPT-5.5 on 23 April, at double the per-token price of GPT-5.4. Gemini 3.5 Flash on 19 May at $1.50 and $9. Opus 4.8 on 28 May. Gemma 4 12B on 3 June, open weights that run on a 16 GB laptop. Claude Fable 5 on 9 June at $10 and $50, suspended worldwide three days later under a US government directive, back about a week after that, and removed from subscription plans in favour of usage credits on 23 June. Gemini 3.5 Pro, which Google said at I/O it looked forward to “rolling it out next month”, has not appeared. Six rate cards, one suspension, one access change and one no-show in ten weeks. Any system with the choice of model written into application code has been wrong, or at least stale, six times since April. This post is the reference design we use to keep that choice out of the code.
Why routing on difficulty fails
The first router most teams build estimates how hard a request is, from prompt length, keyword rules or a small classifier trained on a few hundred labelled examples, and sends the hard ones to the expensive model. It fails for three reasons.
Difficulty is not observable before the answer. “Is this invoice from a supplier we have paid before?” is nine words and may require reconciling three systems. “Summarise this 40-page contract” is long and, for a current model, routine. Length and vocabulary say almost nothing about how much reasoning a task needs, so a length-based router is a coin toss with a cost attached.
Difficulty is one dimension; the decision has at least four. Whether a request may leave the building, how fast it must return, what quality is acceptable and what it may cost are separate constraints. A difficulty score cannot encode that a supplier email containing bank details must never reach a cloud provider, or that a classification step has five seconds and a reply draft has four hours. Those facts are known by the person who wrote the workflow, not by a classifier looking at the prompt.
A predicted route cannot be audited. When finance asks why the bill rose, or a regulator asks why a document went to a particular provider, “the router thought it looked hard” is not an answer anyone can act on. Declarations can be reviewed in a pull request. Predictions drift with the traffic.
Routing on a declared task class fixes all three. The application says what it is doing, not how hard it thinks the instance is: classify.structured, draft.external, escalate.memo. The class is a key into a registry that carries every constraint the router needs. The developer who writes the step knows its class in the way they know which table they write to. Which model serves it this week is not their concern.
The task-class registry
The registry is a versioned configuration file, reviewed like code, with one entry per class and seven mandatory fields.
Latency budget. A p95 target for the whole call, and whether a person is waiting. It decides more routes than any other field. Artificial Analysis measures Opus 4.8 at max effort at about 51 seconds to first token and GPT-5.5 at xhigh at about 44; Gemini 3.5 Flash at its high setting at about 16.5; GPT-5.5 Instant at 0.9. A class with a five-second budget cannot be served by a frontier model at high effort at any price, and the registry should make that impossible.
Quality bar. The name of the evaluation set and the threshold on it: F1 above 0.94 on the triage set, or a pairwise win rate against the current tier on 200 held-out drafts. The bar is what lets a cheaper model take a class: without it every migration is an argument, with it a bench run.
Data sensitivity. Public, internal or restricted. Restricted classes may only use tiers inside the organisation’s own boundary, which since 3 June can mean Gemma 4 12B on a single 16 GB card rather than a rack of them. The router enforces this before it looks at price, and a restricted class’s fallback chain never contains a cloud tier, however busy the local one is.
Allowed tiers, in order. Tiers, not models. A tier is a named level of capability and price, say fast-cheap, long-context, reasoning and flagship; the tier-to-model mapping lives in a separate file. The order is the fallback chain. The first tier that passes the bar within budget is where the class runs; the rest are where the breaker moves it.
Effort pin. The reasoning setting the class runs at, fixed per class and never inherited from a provider default. It is the field most often left blank and most often behind an unexplained bill.
Endpoint. Standard, flex or batch. A class with a four-hour budget has no business on the standard endpoint, and the router, not the engineer who copied the SDK example, should know it.
Cache prefix. The identifier of the stable block the class’s prompts begin with, so assembly and the ledger can reason about hit rates. An owner and a change date complete the entry.
Configuration as model
The rule that follows is simple to state and hard to hold: no model name appears in application code. The code calls the router with a class and a payload; the router resolves class to tier, tier to model, model to adapter and endpoint, and returns the result with a cost record attached. The mapping is the only place a string like gpt-5.5 or claude-opus-4-8 exists, and it is versioned.
This is what makes a price change a configuration change. On 23 April the per-token price of OpenAI’s flagship line went from GPT-5.4’s $2.50 and $15 to GPT-5.5’s $5 and $30. With model names in code, that was a hundred pull requests or a hundred quietly doubled line items. With a registry, it was a bench run on the reasoning-tier classes, a decision per class on whether GPT-5.5’s fewer tokens per task made up for the rate, and one commit. When Gemini 3.5 Flash arrived at $1.50 and $9 with a million-token window and an output speed Google puts at four times other frontier models, it was benched against the fast-cheap and long-context classes and promoted where it passed. When Fable 5 arrived at twice Opus 4.8’s price, it went on the flagship allowed list for two classes and nowhere else.
The corollary for procurement: contract for tiers, not models, because a commitment in one vendor’s tokens bets that their rate card stays best for every class for the life of the contract, and the last ten weeks argue against it.
Figure 1. The router as we build it. The application declares a task class and never a model. The registry supplies the constraints, the policy maps tiers to models from versioned configuration, prompt assembly puts the stable prefix first, adapters choose the endpoint, and every call writes a priced record to the ledger. Bench results flow back into policy, which is how a release day becomes a config change.
Effort is a property of the class
Every frontier model exposes a reasoning-effort control, and the providers keep changing what the levels mean. Opus 4.7 added an xhigh level between high and max in April and raised Claude Code’s default to it; Opus 4.8 kept high as the API default and recommends its extra setting for difficult and long-running work; GPT-5.5 runs from low to xhigh. To a router, the same model at two settings is two models with different latency, token use and quality, and a price comparison that does not say where the dial was set is not one.
So the registry pins effort per class. classify.structured runs at the lowest setting its tier exposes, because its bar is met there and its budget permits nothing else. reason.reconcile runs at medium. escalate.memo runs at max, because a human will read the output and the budget is five minutes. When a provider changes a default, nothing in production moves, because nothing was using the default. When it adds a level, the bench says whether any class should move.
There is a cache reason too: on Anthropic’s API a change to the effort value invalidates the cached messages prefix, so a system that varies effort call by call misses its cache call by call.
Cache-aware prompt assembly
One decision in this design pays for the rest: the order of the prompt. All three major providers charge a tenth of the input rate for tokens that hit their prompt cache: $0.50 against $5 on GPT-5.5 and on Opus 4.7 and 4.8, $1 against $10 on Fable 5, $0.15 against $1.50 on Gemini 3.5 Flash, which also charges $1 per million tokens per hour to keep the cache alive. A hit requires the start of the prompt to be identical to the start of a recent one. Anthropic documents the hierarchy: tools, then system, then messages, a change at any level invalidating everything after it; writes at 1.25 times the input rate for a five-minute lifetime or twice for an hour; and a minimum cacheable length of 2,048 tokens on Opus 4.7, 1,024 on Opus 4.8, 4,096 on Haiku 4.5 and 512 on Fable 5. OpenAI’s guidance is the same in shape: stable instructions and shared reference first, changing content last.
The router owns prompt assembly. Every prompt is built as system instructions, then the class’s reference block, then tools in a fixed order, then any per-tenant material, then the conversation so far, then the current task, last. Nothing variable is permitted above the reference block: no timestamp, request identifier, user name or hash-ordered tool list. The block is registered by prefix identifier so the ledger can report the hit rate per class.
Why this dominates the input bill is arithmetic. The reconciliation step below sends GPT-5.5 12,000 input tokens per call, of which 8,000 is a contract-and-policy block shared across every email from the same supplier. Assembled task-first, every token is fresh and the input costs 6 cents a call. Assembled prefix-first with the block warm, it costs 2.4 cents. Same model, same rate card, same answer, 60 percent off the input line, which is most of the bill for anything that reads more than it writes. For an agent that re-reads a long context on every step the ratio is larger still.
Batch and flex for latency-tolerant classes
Every provider now sells the same tokens at half price to anyone who can wait. OpenAI’s batch and flex endpoints run GPT-5.5 at $2.50 and $15. Anthropic’s Batch API runs Opus 4.8 at $2.50 and $12.50, Fable 5 at $5 and $25 and Haiku 4.5 at $0.50 and $2.50. Google’s batch and flex rates for Gemini 3.5 Flash are $0.75 and $4.50. The discount costs nothing in quality; you give up only the clock.
Most back-office work can give up the clock, and the registry records which. A reply a buyer reviews tomorrow morning does not need an answer in four seconds; a nightly evaluation sample does not need one in four hours. The endpoint field sends those classes to batch or flex without the application knowing, and the latency budget keeps the choice safe.
Flex has a failure mode the router must own: OpenAI’s flex processing returns a resource-unavailable error when capacity is short, uncharged, and the SDK’s default timeout is ten minutes. So a flex route carries a deadline: retry with backoff until half the budget is spent, then move to the standard endpoint at full price and record the fallback. A class that falls back often has the wrong budget, and the ledger is how you find out.
Fallbacks and circuit breakers across providers
Two events this month are the argument for the fallback chain. On 12 June, three days after launch, Anthropic suspended Fable 5 worldwide under a US government directive; it came back about a week later behind new access controls. On 19 May Google said Gemini 3.5 Pro would roll out the following month; it has not. A router whose flagship tier had one entry lost its flagship for a week in the first case and never had one in the second.
The allowed-tiers list is the chain. Each tier entry carries a circuit breaker: an error-rate threshold, a latency threshold measured against the class’s budget, and a cool-down. When a breaker trips, the class moves to the next tier in its own list: a restricted class to another local model or to failing closed, a reasoning class from one provider’s medium-effort model to another’s, and never to a tier that has not passed its quality bar. Breakers are per class and per provider, not global, because a provider slow for 100,000-token prompts may be fine for 2,000-token ones.
Every fallback is a ledger event with the original tier, the tier used and the reason, so on the day Fable 5 went dark we could say within an hour which classes had moved and what a month of it would cost. That is the difference between an incident and a line in a weekly review.
The ledger and its three numbers
Every call through the router writes one record: task identifier, class, tenant, provider, model, endpoint, effort, tokens by type (fresh input, cached input, cache write, output including reasoning), latency, outcome (passed, failed, fell back) and the rate-card version used to price it. Pricing happens at write time, against the rate card in force that day, because a rate card is now a time series.
From that record the ledger reports three numbers. First, cost per completed task, by class and tenant, with yesterday beside it. This is the number that detects a tokenizer change, which no rate card shows: Anthropic’s documentation puts its 4.7-and-later tokenizer at about 30 percent more tokens than its predecessor for the same text, so a class moved from Sonnet 4.6 to Opus 4.8 pays for more tokens at an unchanged rate. Second, the cached share of input tokens, by class. When it falls, someone put a timestamp above the reference block. Third, the flagship share: the fraction of tasks, and of spend, that reached the top tier. When it rises without a change to the registry, a breaker is tripping or an escalation rule is loose. Everything else in the ledger is for investigation; these three are for the people who sign the bill.
A worked example: procurement email triage
One task through the router, with the arithmetic. Inbound emails to a procurement shared mailbox: invoice queries, purchase-order changes, quotations, disputes and the occasional request to change a supplier’s bank details, a common form of payment fraud and the reason the first step is where it is. Six steps. Token counts are our estimates; prices are the 25 June list rates cited above, with cached input at a tenth of the fresh rate and reasoning tokens billed as output.
Step
Task class
Tier chosen on 25 June
Why
Approx. cost per 1,000 runs
1. Screen and redact detect bank-detail changes, mask account numbers and personal data
screen.restricted
Local: Gemma 4 12B on a 16 GB GPU
Restricted data never leaves the boundary. Two-second budget. The bar is recall on bank-detail patterns, not fluency.
GPU-seconds at an internal rate; no token bill
2. Classify and extract intent, supplier, PO number, amounts, to a fixed schema
Five-second budget. Passes the F1 bar at $1.50 and $9. Taxonomy, schema and examples sit in the cached prefix at $0.15.
≈ $2 (600 fresh, 1,200 cached, 80 out)
3. Reconcile invoice against PO lines and contract terms; list discrepancies; recommend an action
reason.reconcile
Reasoning: GPT-5.5, effort pinned medium, standard endpoint
Sixty-second budget, money at stake. Cheapest tier that passed the bar. The 8,000-token supplier block is cached across that supplier’s emails.
≈ $84 (4,000 fresh, 8,000 cached, 2,000 out incl. reasoning)
4. Draft the reply for a buyer to review next morning
draft.external
Reasoning tier via flex: GPT-5.5 at $2.50 and $15, effort low
Four-hour budget. Flex is half price; the router falls back to standard if flex has not returned by hour two.
≈ $14 (3,000 fresh, 1,500 cached, 400 out)
5. Escalation memo disputes and bank-detail changes, about 8% of runs
escalate.memo
Flagship: Opus 4.8, effort max; Fable 5 allowed above a value threshold
Five-minute budget. A person reads every one, so the quality bar is the highest in the registry.
≈ $14 (8% of runs; 15,000 in, 4,000 out incl. thinking)
6. Quality sample judge steps 2 to 4 for the bench, 5% of runs
eval.judge
Flagship via Batch API: Opus 4.8 at $2.50 and $12.50, effort high
Twenty-four-hour budget. Feeds the quality bars that every other row depends on.
≈ $1 (5% of runs; 6,000 in, 300 out)
Total per 1,000 triaged emails
Plus local GPU time for step 1.
≈ $115
Table 1. One procurement email through six task classes at 25 June 2026 list prices: Gemini 3.5 Flash $1.50 and $9 (cached $0.15), GPT-5.5 $5 and $30 (cached $0.50, flex $2.50 and $15), Opus 4.8 $5 and $25 (batch $2.50 and $12.50). Token counts are estimates; the arithmetic is the point. Step 3 is three-quarters of the cloud cost and is the row the bench re-tests on every release.
Three things to notice. Reconciliation is about three-quarters of the cloud cost, and it is on GPT-5.5 at medium effort not because that is the best model but because it is the cheapest tier that passed the class’s bar on our bench. Gemini 3.5 Flash at its high setting would run the step for about $25 per 1,000 and is re-benched on every release; the day it passes, the registry changes and nothing else does. Escalation uses Opus 4.8 at max effort on 8 percent of runs; Fable 5 is on the allowed list and would cost about $28 per 1,000, double, a fair price above a value threshold and a poor one for everything. And the restricted step never produces a token bill; the ledger carries it in GPU-seconds at an internal rate, the right unit for a model that costs the same busy or idle.
Run the same six steps on one flagship, Opus 4.8 at default effort on the standard endpoint with task-first prompts so nothing caches, and the same token counts come to about $170 per 1,000 runs, roughly half as much again. More to the point, it cannot meet the two-second budget on step one or the five-second budget on step two, and it sends bank details to a cloud provider. The saving is the smaller reason to route.
Four anti-patterns
Model names in code. The commonest and, over time, the most expensive. Every one is a pull request on every price change, and the ones nobody finds pay April’s rates in December. The test is mechanical: a search of the application repositories for any provider’s model identifier should return only the router’s configuration.
Routing by prompt length. It sends the nine-word fraud question to the cheap tier and the 40-page summary to the expensive one, backwards, and says nothing about latency or sensitivity. If a team has it, it is usually because the registry does not exist yet.
One effort setting for everything. Usually the provider’s default, usually high, usually inherited without anyone deciding. It makes the fast classes slow and the cheap classes expensive, and when the provider changes the default the bill changes without a commit.
Caches defeated by rewriting history. Agents that compact or summarise their conversation each turn change the prefix each turn and miss the cache on every token after the system block. Append-only histories cache; rewritten ones do not. The same applies to any variable material placed early: a date in the system prompt, a session identifier, a tool list in hash order. Each is one line, and each costs a tenfold difference on most of the input bill.
Making the next price change boring
None of this is novel or large; the router is a few thousand lines and the registry is a file. What it buys is that the next release, and there will be one next week, is a config change with a bench run attached rather than a migration. In build order:
Write the registry before the router. Listing every task class with its latency budget, quality bar and sensitivity is most of the work, and it surfaces the classes nobody owns.
Remove model names from code and put the mapping in one versioned file. Make the search for provider identifiers part of continuous integration.
Pin effort per class and bench on release day. Treat a provider’s default as a suggestion you declined.
Let the router assemble prompts, prefix first. Register the stable block per class and report the hit rate.
Route latency-tolerant classes to batch or flex, with a deadline and a fallback. Half price for the clock is the best trade on any rate card.
Give every class a fallback chain that respects its sensitivity. Breakers per class and per provider, and every fallback a ledger event.
Price at write time against a dated rate card, and report three numbers. Cost per completed task, cached share of input, flagship share.
GPT-5.5 Instant is a good model, and I expect several of our fast classes will move to it once the bench has run. Nobody in the application teams will need to know.
Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn