Built for the Real World · Essay · Unit economics

The AI bill nobody owns: a chargeback playbook from the platform seat

GPT-5.6 arrived in three tiers on 9 July and two were cut on 30 July; Kimi K3 charges a cache hit a tenth of a miss; Opus 5 kept the old Opus price. A bill that moves like that cannot be allocated from an invoice. Here is how we tag every call at the gateway, price it the day it happens, and show it before charging it.

Abstract illustration for “The AI bill nobody owns: a chargeback playbook from the platform seat”

This week finance asked me a question I could not answer from an invoice: which business unit caused July’s AI bill. The bill arrives as one line from each of four providers and one from a GPU pool, and none of those lines knows that a loan was processed or a ticket closed. That is the bill nobody owns. On 9 July OpenAI made GPT-5.6 generally available as three tiers: Sol at $5 and $30 per million input and output tokens, Terra at $2.50 and $15, Luna at $1 and $6. On 17 July Moonshot announced Kimi K3 with a price list on which the same input token costs $0.30 if it hits the cache and $3 if it does not. On 24 July Anthropic shipped Claude Opus 5 at $5 and $25, the price Opus 4.8 carried. On 30 July OpenAI cut Luna by 80 percent to $0.20 and $1.20 and Terra by 20 percent to $2 and $12. On 11 August NVIDIA released Nemotron 3.5 Lightning, 31.6 billion parameters with 3.6 billion active; on 12 August Alibaba published the weights of Qwen3.8-Max, 2.4 trillion parameters. Last week Anthropic cancelled a scheduled 1 September increase for Sonnet 5 and made $2 and $10 the standard price.

Six price events in six weeks, two of them for models with no price per token at all. A bill that moves like that cannot be allocated once a quarter from a spreadsheet. This is the playbook we use to attribute every call to the business unit that caused it, from the platform team’s seat.

Why an AI bill is not a cloud bill

The State of FinOps 2026 survey, 1,192 practitioners representing more than $83 billion of annual cloud spend, found that 98 percent now manage AI spend, up from 63 percent in 2025 and 31 percent in 2024, and that what they find hardest is, in order, seeing the AI costs at all, allocating them to business units, and working out what they return. Those are the first three steps of cloud FinOps, done once already by the same people, and they are hard again because four properties of AI spend break the assumptions cloud allocation rests on.

Tokens, not hours. A virtual machine costs the same per hour busy or idle, so you allocate it by who owns it, from the tag, a week later. A model call costs by what went in and came out, so you allocate it by who made the call, the only moment at which that is known.

Prices that change weekly. A cloud rate card changes a few times a year. In six weeks a Luna token fell 80 percent, a Terra token 20 percent, and Opus 5 arrived at an unchanged price with a different token count per task, because a new generation spends tokens differently on the same work. A forecast that holds the rate card constant is wrong by the end of the quarter.

Caches that change the price of the same request tenfold. On Kimi K3 a cached input token is $0.30 and a fresh one $3. On Sol the pair is $0.50 and $5; on Opus 5, $0.50 and $5 with a cache write at $6.25; on Luna, $0.02 and $0.20. The same 10,000-token prefix costs five cents on Sol fresh and half a cent cached, depending on whether an engineer put a timestamp at the top of the prompt. No cloud bill has a line whose price is set by prompt hygiene.

Agents that spawn their own usage. A chat request is one call. An agent is a loop that calls the model once per step with the whole history, retries failed steps, and increasingly starts sub-agents that do the same. The number of calls per task is decided at run time by the model, not by the developer, so the driver of the bill is a quantity nobody set, and attribution has to happen inside the loop rather than from the invoice.

The Foundation’s guidance has moved with this. Its FinOps for AI overview of 17 February treats tokens as a new cost unit, recommends tagging by project, environment, workload, team, cost centre and purpose, and describes showback as giving teams sight of their AI costs “without immediately charging them”.

The unit of accountability: task class and tenant

Every call that reaches a model through our platform carries two tags, and the gateway refuses a call that lacks either. The first is the task class, the declared purpose of the call, which I described in June as the key the router uses to pick a tier: classify.structured, draft.external, escalate.memo. The second is the tenant, the business unit or cost centre that will see the cost on a statement. Both are set at the call site, not inferred from the prompt, and validated against a registry.

The class says what the money bought, and therefore which tier, effort setting and cache prefix were appropriate; the tenant says who should see it. A record with a tenant and no class tells finance who spent but not whether the spend was reasonable; one with a class and no tenant tells the platform what the work cost but not whom to show it to.

Enforcement is the part teams skip. The FinOps Framework’s allocation capability describes the mature state as one in which metadata fields are a precondition for provisioning and requests that lack them are blocked. For AI the resource being provisioned is the call itself, so the gateway returns an error, naming the missing tag, for any request without both. We switched this on in the second month of showback, for reasons I will come back to. It follows that no API key exists without a tenant and a named owner: a key is a principal, in the sense I argued for agents in July, and an unowned key is spend with nobody to show it to.

Agents need one more rule: a sub-agent inherits the tenant of the task that started it and declares its own class, because the unit that asked for the task is the one that benefits, whichever team wrote the sub-agent. The alternative, a shared agent service with its own tenant, produces a cost centre called agents that grows every month and belongs to nobody.

The ledger record and the versioned rate card

The gateway writes one record per call: task identifier, class, tenant, key owner, provider, model, endpoint, effort, tokens by type (fresh input, cached input, cache write, output including reasoning), GPU-seconds if the model is self-hosted, latency, outcome, and the version of the rate card used to price it. The last field makes the record a ledger rather than a log: pricing happens at write time, against the rate card in force that day, so the cost never changes, and the rate card is a versioned table with effective dates.

This summer the table received rows on 9 July for the three GPT-5.6 tiers, 17 July for Kimi K3, 24 July for Opus 5, 30 July for the Luna and Terra cuts, and last week for Sonnet 5’s permanent price. Each row has a start date and, where the provider announced one, an end date, because a promotional rate with an end date is a liability the forecast must find. Maintaining the table is a few hours a week in a month like July, and those hours make every other number in this post trustworthy.

Keep the ledger separate from the providers’ dashboards, because the ledger is the only place where a Luna call, an Opus 5 call and a Nemotron call on our own GPUs appear in the same units. The evaluation bench reads the same ledger, so a change of tier for a class shows up as a cost change the next day rather than at month end.

The attribution path from application to showback statementBlock diagram: an application calls the gateway with a task class and a tenant; the gateway admits or rejects the call, checks the tenant’s budget, routes through provider adapters, meters tokens and GPU-seconds, and writes a record to the ledger, which is priced against a versioned rate card and summarised into a monthly showback statement; spend to date feeds back into the budget check. GATEWAY Application run(class, payload) tenant from identity Tag and admit class · tenant · key owner untagged calls rejected Budget check alert · degrade tier · stop per tenant, per class Provider adapters OpenAI · Anthropic · Moonshot self-hosted · GPU-seconds Meter tokens by type · GPU-seconds latency · outcome Rate card versioned, dated rows 9 Jul · 24 Jul · 30 Jul Ledger one record per call priced at write time Showback per tenant, per class then chargeback class record price of the day spend to date A call without a class and a tenant does not run. Every call that runs is priced the day it happens, against the rate card of that day.
Figure 1. The attribution path. The application declares a task class and carries a tenant from its identity; the gateway admits or rejects the call, checks the tenant’s budget, routes it, meters tokens or GPU-seconds, and writes one priced record to the ledger. The rate card is a dated table, and the ledger’s spend to date feeds the budget check that can stop the next call.

Showback first, chargeback second

The Foundation is careful here: its invoicing and chargeback capability says that “Showback is always required in any FinOps practice” while chargeback depends on accounting policy, and that neither is more mature than the other. Kong, which sells a gateway that meters this, calls showback “visibility without consequence” and chargeback the version with consequences. FinOps LLM, a practitioner site, recommends showback first so that teams can clean their tags and build dashboards “before anyone’s budget is on the line”. The order matters more than the destination.

We ran six months of showback before any money moved. Month one: a statement per tenant, by task class, with the unattributed share as a line of its own, sent to each unit’s engineering lead, not to finance. Month two: the gateway starts rejecting untagged calls, and the unattributed line falls to near zero because the alternative is that the application stops. Month three: the statement gains a forecast built from the daily run rate and the rate card with its announced changes, and the lead signs it. Month four: the unit chooses its outcome unit, the thing it counts, and the statement reports cost per outcome beside cost per task. Month five: a shadow chargeback, the journal entries that would have posted, shown to finance and the unit together, with disagreements settled as tagging corrections. Month six: the entries post.

There is no budget fight because by month six nobody learns anything from the invoice: every number has been on a statement for months, and the arguments have been had over tags, which are cheap to change, rather than over money. The fight people expect comes from the first statement being the one with consequences.

Budgets that stop the transaction

A statement that arrives on the fifth of the next month is a postmortem. The control that prevents the incident is a budget the gateway checks before the call, which is possible only because the ledger is priced at write time.

Each tenant carries a monthly budget per task class, agreed during showback, and the gateway holds three thresholds against it. At the first, typically 80 percent of the budget prorated to the day, the owner is told and nothing changes. At the second the gateway degrades: latency-tolerant classes move to the batch endpoint at half price, flagship classes fall to the next tier that passed the bench, and effort pins drop a level where the class allows it. At the third, discretionary classes stop and the call returns an error that says why; customer-facing classes continue and page the owner. The business unit decides in advance which classes are discretionary, not the platform during the incident.

Agents get a budget per task on top, denominated in money rather than steps, because a step on Sol and a step on Luna differ in price by a factor of twenty-five. A loop that may take unlimited steps will, on hard cases, take them; a budget with an explicit action to request more makes the cost bounded. Agents that give up early to stay under budget are a separate failure, caught by the success rate beside the cost line.

The three numbers each unit gets, and the fourth number finance wants

The statement reports three numbers per tenant and per class, the same three I gave the router in June. Cost per completed task, with last month beside it, which detects a generation that spends more tokens on the same work. Cached share of input, which falls the day someone puts a date above the reference block and is the number most often behind a bill that rose while the price fell. Flagship share, the fraction of tasks and of spend that reached the top tier, which rises when an escalation rule is loose or a breaker is tripping. A platform team can compute all three without knowing anything about the business.

Finance wants a fourth, the only one that can justify the spend rather than explain it: cost per business outcome, per loan processed, per ticket resolved, per invoice reconciled. The Foundation’s May 2025 paper on AI business value asks for cost mapped to each use case and tracked on a unit basis such as cost per interaction, and its token-economics piece this May names business-unit showback or chargeback as the governance step that connects token consumption to value. The platform cannot compute this number alone; the unit has to declare its outcome and send the identifier on the call, or join its own records to the ledger afterwards. Month four exists for this.

It is also the number that makes a model change discussable with someone who does not know what a token is. A lending unit that processes an application for four and a half cents of model time does not care whether reconciliation moved from Terra to Kimi K3; it cares whether the number moved and whether approvals held. Pair it with the success rate and the escalation rate, as the tokens-to-done bench does, or a cheap failure will look like a win.

Self-hosted: GPU-seconds as a currency

August put self-hosting on the same statement. Nemotron 3.5 Lightning, released on 11 August under the OpenMDW licence, is 31.6 billion parameters with 3.6 billion active, produces about 670 tokens per second on a hosted endpoint and fits on one GPU. Qwen3.8-Max, whose weights Alibaba published on 12 August, is 2.4 trillion parameters with 95 billion active and needs a cluster; Kimi K3’s weights, published on 27 July, are about 1.4 terabytes that Moonshot says want 64 or more accelerators. I covered what sovereignty costs on 14 August. The narrower question here is how a model with no price per token appears on a business unit’s statement.

Treat GPU-seconds as a currency and the ledger as the exchange. Every call to a self-hosted model records the GPU-seconds it consumed, and the ledger prices them at an internal rate, the pool’s monthly cost divided by its GPU-hours. A busy second is charged to the tenant whose call used it. The trap is the idle second. A pool costs the same serving or waiting, and the Foundation’s paper on generative AI usage reports GPU resources often running at 15 to 30 percent of capacity. Spread idle cost across tenants in proportion to usage and the busiest tenant pays most for capacity nobody used; fold it into the internal rate and the rate is wrong by a factor of three and the hosted comparison is meaningless.

So idle cost is a line of its own, charged to the platform and printed on every statement, and it decides whether the pool is the right size and whether self-hosting is the right choice at all. DeepInfra serves Nemotron 3.5 Lightning at $0.05 and $0.20 per million tokens; a two-GPU pool at 60 percent utilisation has to beat that on busy seconds alone, or the restricted-data argument has to carry it. The token-economics piece puts the general case plainly: self-hosted inference carries the highest capital commitment and “the lowest marginal cost per token at sustained scale”, and the word that matters is sustained.

What went wrong when we did it

Shared keys with no owner. The first statement, in the spring, had an unattributed line of roughly a fifth of hosted spend, most of it from three API keys created for pilots in 2025 and passed between teams since. Nobody would claim them, because claiming them meant owning the bill. The fix was not a tagging campaign but the rejection rule: once a call with no tenant did not run, the keys acquired owners within a week because the applications behind them stopped.

Agents charging the wrong tenant. A shared customer-care agent, built by the platform team and used by two business units, carried the platform’s own tenant on every sub-agent call, because the framework took the tenant from the service identity rather than the task. For two months the platform was the group’s third-largest consumer of Terra tokens and could not say why. The fix was the inheritance rule above, enforced by a gateway test.

The month the price fell and the bill rose. On 30 July Luna’s price fell 80 percent. In the first two weeks of August the customer-care unit’s Luna line went up. A team removed the step cap from its conversation-summary agent, reasoning that the tier was now close to free, and call volume on the class quadrupled. A deployment on 4 August then added the conversation’s start time to the system prompt, above the cached block, and the class’s cached share fell from above 80 percent to under 30. A price cut, fourfold volume and a cache rate down by two-thirds add up to a higher bill, and the ledger showed all three the next morning because it reports volume and cached share beside price. A price cut is a re-forecast, not a saving, until the ledger confirms the volume held.

A worked example: three business units, one month

Three business units of a group, one month, at the list prices in force on 17 August. Volumes and tokens per task are illustrative, ours and rounded; the prices, listed under the table, are the providers’. The self-hosted pool is two GPUs running Nemotron 3.5 Lightning for restricted-data classes: 1,440 GPU-hours at an illustrative internal rate of $3 per GPU-hour, $4,320 for the month, 60 percent utilised, so $2,592 of busy seconds and $1,728 idle.

Business unitTask classes and tiers on 17 AugustTokens, millions: fresh / cached / outputHosted costSelf-hosted, GPU-secondsTotalCost per outcome
Consumer lending
120,000 applications (illustrative)
extract and classify on Luna; reconcile on Terra; escalate on Opus 5 at high effort (8% of applications); PII screen on the pool888 / 2,904 / 276$3,936$1,426 (55% of busy)$5,3624.5¢ per application
Customer care
600,000 conversations (illustrative)
intent, routing and summaries on Luna; reply drafting on Terra; complaints on Opus 5 (3%); redaction on the pool3,354 / 12,780 / 954$9,102$907 (35% of busy)$10,0091.7¢ per conversation
Procurement
60,000 invoices (illustrative)
screen on Luna; contract reconciliation on Kimi K3 with a one-million-token context; nightly quality judge on Opus 5 batch (5%); bank-detail screen on the pool450 / 2,280 / 99$1,493$259 (10% of busy)$1,7522.9¢ per invoice
Platform
not charged to any unit
idle GPU time in the pool; evaluation bench runs on Opus 5 batch40 / 0 / 4$150$1,728 idle$1,878not applicable
Total
one month, list prices of 17 August
Luna 69% of hosted tokens and 13% of hosted cost; Opus 5 and Kimi K3 together 34% of hosted cost4,732 / 17,964 / 1,333$14,681$4,320≈ $19,000
Table 1. A monthly showback statement for three business units at the list prices of 17 August 2026: GPT-5.6 Luna $0.20, $0.02 cached and $1.20; Terra $2, $0.20 and $12; Claude Opus 5 $5, $0.50 and $25, batch $2.50 and $12.50; Kimi K3 $3, $0.30 and $15. Volumes and tokens per task are illustrative; the GPU pool is priced at an illustrative internal rate of $3 per GPU-hour. The platform row is the only one without a business owner, by design.

Four things to read off the statement. Customer care is more than half of the bill and the cheapest unit of work on it, at under two cents a conversation; volume, not extravagance, is what finance is looking at. Lending’s escalation class, 8 percent of applications on Opus 5 at high effort, is nearly half of its hosted cost and the row the bench re-tests on every release. Procurement’s Kimi K3 line is 90 percent cached; the same tokens fresh would cost about $4,100 rather than $1,224, which is why cached share sits next to cost. And the platform row, $1,878 of idle GPU time and bench runs, has no business owner by design: it should be visible, argued over and small.

Run the same month at the 9 July prices and customer care’s Luna line is $6,960 rather than $1,392. Nothing about the work changed. That is what a moving rate card does to a budget set in June, and why the gateway’s budget is re-based the day the rate card changes.

Recommendations

  1. Tag at the gateway or not at all. Task class and tenant on every call, validated against a registry; untagged calls rejected.
  2. No key without an owner and a tenant. Retire the ones nobody will claim.
  3. Sub-agents inherit the tenant of the task that started them. Test for it in the gateway.
  4. Price at write time against a versioned rate card with effective and end dates. Maintain the table as a job with a name on it.
  5. Six months of showback before a journal entry posts. Statements first, enforcement second, forecasts third, outcome units fourth, a shadow chargeback fifth.
  6. Budgets the gateway checks before the call. Alert, degrade, stop, per tenant and class; the unit decides in advance which classes may stop.
  7. Three numbers per class and one per unit. Cost per task, cached share and flagship share for engineering; cost per outcome, beside the success rate, for finance.
  8. Meter self-hosted models in GPU-seconds. Charge busy seconds to tenants, print idle seconds as a platform line, never spread them.
  9. Treat a price cut as a re-forecast. Re-base budgets the day the rate card changes and watch volume and cached share for a fortnight.

The bill nobody owns is not a finance problem. It is a missing field on a call. Add the field, refuse the call without it, price it the day it happens, and six months later the chargeback is a formality.

Sources

  1. OpenAI. Sol, Terra, and Luna, our GPT-5.6 family of models, are starting to roll out now in ChatGPT, Codex, and the API. 9 July 2026. Eden AI, OpenAI cuts GPT-5.6 API prices: Luna falls 80%, Terra 20%, Sol holds. 30 July 2026.
  2. Anthropic. Introducing Claude Opus 5. 24 July 2026. Cache, batch and Sonnet 5 pricing from Claude Platform pricing.
  3. Constellation Research. Moonshot AI launches Kimi K3. 16 July 2026.
  4. Artificial Analysis. NVIDIA launches Nemotron 3.5 Lightning. 11 August 2026. Hosted price from DeepInfra, NVIDIA Nemotron 3.5 Lightning release, 11 August 2026.
  5. LLM Stats. Qwen3.8-Max open weights: first Max-class Qwen you can download. 12 August 2026.
  6. FinOps Foundation. FinOps for AI Overview, 17 February 2026; Invoicing & Chargeback and Allocation, FinOps Framework; State of FinOps 2026.
  7. FinOps Foundation. Unlocking AI Business Value with FinOps, 13 May 2025; Optimizing GenAI Usage: A FinOps Perspective on Cost, Performance, and Efficiency, 23 May 2025; J.R. Storment, Token economics: the atomic unit of AI value, 10 May 2026.
  8. Kong. LLM cost management: AI showback and chargeback. 6 April 2026. FinOps LLM, Chargeback vs showback: picking the right FinOps model, updated 1 August 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn