On 5 October, Anthropic announced that Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5 can now be called through an India endpoint on Amazon Bedrock, with every request processed inside the country. The plumbing had gone live on 29 September, when AWS published three India inference profiles, in.anthropic.claude-opus-5, in.anthropic.claude-sonnet-5 and in.anthropic.claude-haiku-4-5, which route each call between its Mumbai (ap-south-1) and Hyderabad (ap-south-2) Regions and nowhere else. Reliance and CRED tested the endpoint in a private preview. Kotak, Axis and IndusInd banks already use Claude, NPCI is building an agentic platform it calls AiNxt on it, TCS is rolling it out to 50,000 associates, Infosys has stood up an Anthropic centre of excellence, and Mahindra, Godrej, Swiggy, Zomato, Freshworks and Razorpay are customers. Irina Ghose, Anthropic’s managing director for India, said the company had “committed to bringing Claude inference to India” in August and was now delivering. Sandeep Dutta, who runs AWS in India and South Asia, supplied the commercial logic: in regulated industries, data residency can decide whether AI deployments “move from pilots into production”.
Seven weeks earlier AWS had done the same for OpenAI, adding India profiles for GPT-5.6 Terra and Luna on 17 August. Residency has become a product feature that the hosted labs sell, and they are shipping it to India within weeks of each other. In August I argued in a post on the trillion-parameter open-weight wave that open weights win on data that cannot leave, because the hosted option exists only where the provider has built a region. That was true in August and is less true in October. The short version of what follows: residency is now cheap to buy and costs you a model generation; and residency, even bought, is not sovereignty.
What shipped, and how the routing works
Bedrock has two kinds of cross-Region inference profile. A global profile lets AWS process a request in any commercial Region in the world and is priced about 10 percent below the alternative. A geographic profile restricts processing to Regions inside one geography, at standard pricing, and exists for compliance. The India profile is the geographic kind: a request sent to the Mumbai endpoint may be served in Mumbai or Hyderabad, depending on capacity, and the routing decision is AWS’s. The documentation is specific about the data. Traffic between the two Regions stays on the AWS network, encrypted in transit; customer data is “not stored in a destination Region”; CloudTrail records every request in the source Region, with an inferenceRegion field that says which of the two served it; Bedrock’s zero-data-retention default applies; Guardrails, invocation logging and the Converse and Messages APIs all work. Pricing is the standard Bedrock rate, charged to the source Region, with no routing surcharge.
The same release added Seoul profiles for Opus 5 and Sonnet 5 and a Singapore profile for Sonnet 5, so India was one of a batch. The India endpoint is not a sovereign cloud, a dedicated cluster or a special legal arrangement. It is a routing constraint on a shared fleet, delivered as a string in the model field, and its whole value is that it is boring.
Anthropic’s part is the model and the go-to-market: a Bengaluru office opened in February under Ghose, a partner network launched in March, a first partner certification programme in Bengaluru in August with a target of 5,000 people, and TCS as global premier partner. The OpenAI profiles, in.openai.gpt-5.6-terra and in.openai.gpt-5.6-luna, use the same two Regions and the same zero-retention model, with one visible difference: AWS’s OpenAI post says flagged content goes to offline abuse detection, and does not say where.
The generation lag
Here is the catch, and it is the most important thing in this post for anyone holding a budget. The India endpoint carries Opus 5, Sonnet 5 and Haiku 4.5. The global endpoint carries Opus 5.5, released on 22 September, Sonnet 5.5, released on 28 September, and Fable 5.1, released on 1 September. Opus 5.5 launched on Bedrock with US, EU, Australia, Japan and global profiles and no India profile; Sonnet 5.5 launched with the global profile only; Fable 5.1 has US and global profiles. So a bank that insists on in-country Claude inference is, on Anthropic’s own benchmark tables, one generation behind the bank next door that does not. I wrote last week about what that generation is worth: Sonnet 5.5 scores 70.6 on Terminal-Bench 4.0 against 66.4 for Opus 5.5, uses fewer tokens for the same work and generates output more than 30 percent faster than Sonnet 5; Opus 5.5 is 20 percent cheaper than Opus 5 on list and 60 percent cheaper on cache reads, $0.20 against $0.50 per million. Those are first-party list prices, not Bedrock’s, but the shape of the gap carries over.
So the price of residency is concrete: roughly 10 percent on rate against the global profile, plus the per-task efficiency, the cache cut and the benchmark points of the generation you cannot have. For an agentic workload whose bill is mostly cache reads, in-country Opus 5 costs about two and a half times what offshore Opus 5.5 costs for the same context. For a chat workload on Sonnet, where the list price did not change, the cost is the up to 30 percent fewer tokens per task that you do not get.
This is not an Anthropic quirk; it is how every hosted lab ships. OpenAI’s India profiles carry GPT-5.6 Terra and Luna; GPT-6 Sol and Luna reached Bedrock on 22 September, and the announcement lists no India profile. Microsoft’s documentation for Azure OpenAI is unusually candid: new models arrive as Global deployments first, then Data Zone, then geography-based, and the geography-based types “have no guaranteed availability date”. Claude on Vertex AI is the starkest case. Anthropic’s own documentation says regional endpoints support Sonnet 4.6 and earlier, that newer models use the global endpoint or the US and EU multi-region endpoints, and that regional and multi-region endpoints carry a 10 percent premium; there is no India multi-region endpoint. Capacity for a new model is scarce at launch, the labs put it where demand is densest, and a two-Region geography on the far side of the world gets it when it gets it. Table 1 sets the offers side by side.
What residency covers, and what it does not
Residency, as sold, covers one thing: where the compute that runs inference sits, and therefore where your prompts and outputs exist in plaintext for the milliseconds that matter. On Bedrock’s India profile that commitment is clean, and so is the logging, because CloudTrail and CloudWatch write in the Region you called. Beyond that ring the picture changes. Figure 1 lays it out as five layers from the inside out.
Prompts and outputs. Covered. Processed in Mumbai or Hyderabad, not stored in the destination Region, zero retention by default.
Logs and telemetry. Partly covered. Your invocation logs and audit trail stay in the source Region. What is not stated, for either lab, is where safety and abuse classification runs when a request is flagged, who reviews it and under which jurisdiction. For most workloads that is a footnote; for a bank’s customer transcripts it is a question to ask in writing.
Control plane. Not covered. Which Regions the profile may route to, the quota in each, the model catalogue, the Guardrails service, the release train that decides when Opus 5.5 gets an India profile: all of it belongs to AWS and Anthropic. Inference profiles do not support Provisioned Throughput, so the capacity you get in India is the capacity AWS chose to put there.
Weights. Not covered, and never will be on a hosted offer. The model file is Anthropic’s, on AWS hardware, in AWS accounts; you cannot inspect it, copy it, fine-tune it in place or pin it against a retirement. This is the layer open weights give you and nothing else does, and the layer most of India’s sovereign-AI conversation is actually about.
Legal entity and terms. Not covered. Your contract is with the cloud provider, and the model reaches you through its marketplace terms; the lab is not your counterparty. Lifecycle is the case in point. Anthropic’s platform documentation says lifecycle dates on partner-operated platforms are “set by the partner and can differ from the Claude API schedule”. The India profile you adopt today has a retirement date you have not been told, set by a company you may have no direct relationship with, and its successor has no India profile yet.
None of this is an argument against buying residency; it is an argument for naming what you bought. In August I listed five layers of sovereignty, weights, licence, compute, data and controls, and said open weights give you one. Hosted residency gives you a different one, the data layer, with the lab’s controls thrown in. They still do not add up to five.
The regulatory map that makes this matter in India
Indian regulation is layered, and the layers point in different directions.
DPDP. The Digital Personal Data Protection Act of 2023 is, on its face, permissive about transfers. Section 16 lets the central government restrict transfers to countries it names, a negative list, and section 16(2) preserves any other law that imposes a higher degree of protection. The Rules notified in November 2025 keep that structure and add the thing that bites: Rule 12(4) lets the government specify categories of personal data that Significant Data Fiduciaries must keep in India, along with the related traffic data, and the largest banks and platforms should expect that designation. The substantive obligations, including 72-hour breach notification to the Data Protection Board and penalties of up to ₹250 crore, take effect in May 2027. DPDP does not mandate localisation of inference; it creates a switch under which localisation of named categories can be turned on, and it leaves every sectoral rule in force.
RBI and payments data. The sectoral rule that matters most is older and sharper. The RBI’s circular of 6 April 2018 required that end-to-end transaction data for payment systems be stored only in India, with compliance by 15 October that year. Its FAQ of 26 June 2019 allowed processing abroad provided the data is deleted from the foreign system and brought back to India “not later than the one business day or 24 hours” after processing, whichever is earlier, and extended the obligation to banks, payment system operators and the third-party vendors they engage. Against a hosted model offshore, every prompt containing a UPI transaction is payments data processed abroad, and you owe the regulator evidence that it was deleted within 24 hours from an inference stack you do not operate. In-country inference does not change the rule; it removes the proof burden, which in practice is the whole difficulty.
RBI on AI itself. On 13 August 2025 the RBI published the report of its FREE-AI committee, chaired by Pushpak Bhattacharyya of IIT Bombay: seven guiding sutras and 26 recommendations across six pillars, among them financial-sector data infrastructure, indigenous models for financial services, an innovation sandbox, a standard AI incident report and clearer liability. It is advisory, but it signals where supervision is going: a bank will be asked not only where its data is but which model made the decision and who answers for it.
Securities and government. SEBI’s cloud framework of 6 March 2023 requires a regulated entity’s data to be stored and processed in data centres of cloud providers empanelled by MeitY, with the entity retaining ownership of its data, logs and keys. For the public sector, which AWS and Anthropic named as a target segment, processing location is the first gate in any procurement.
Nothing in Indian law says a large language model must run in India. Several things say particular data must stay in India, or come back within a day, or be kept under conditions the government can tighten by notification. Residency as a product is valuable because it collapses a dozen attestation problems into one configuration value a regulator can read in a CloudTrail log.
Three options for a regulated Indian enterprise
The choice for any given workload comes down to three postures, compared on cost, capability, control and compliance.
Open weights on your own GPUs. Cost: hardware, two to three platform engineers and, for a trillion-scale model such as Kimi K3, a six-figure monthly infrastructure line before a token is served; for a small model such as Nemotron 3.5 Lightning, a single GPU. Capability: a few benchmark points behind the hosted frontier at the top, level with it on bounded tasks. Control: complete, including the whole job of containment. Compliance: trivially satisfied on location, but the licence has to clear, and Qwen3.8-Max shipped without one.
In-country hosted frontier. Cost: standard Bedrock rates, roughly 10 percent above the global profile, on a model a generation back. Capability: Opus 5, Sonnet 5 and Haiku 4.5, or GPT-5.6 Terra and Luna. Control: the lab’s controls as a service, none of the weights, a lifecycle you do not set. Compliance: processing location, retention and logging are clean and evidenced; the classifier and support questions need a written answer.
Offshore hosted frontier. Cost: the cheapest per task, with the newest cache pricing. Capability: Opus 5.5, Sonnet 5.5, Fable 5.1, GPT-6 Sol and Luna, and whatever ships next week, including the gated tiers no in-country endpoint will see for a long time. Control: as above. Compliance: only for data classes that may leave, a determination your data classification must be able to make per request, not per project.
Notice what the in-country option did to the open-weights case. In August the argument for self-hosting a 2.4-trillion-parameter model was that for regulated data there was no alternative. Now there is one, for most regulated data, at a fraction of the operational effort. Open weights retreat to where they were always strongest: the bounded, high-volume tier where a 3.6-billion-active-parameter model on one GPU beats any API on price, and the workloads where the weights themselves are the requirement, because you need to fine-tune, pin a version for seven years or run air-gapped. Smaller territory than eight weeks ago, and still real.
A worked example: one bank, three workloads
Take a mid-sized private bank with three things on the AI roadmap. The token counts are ours and illustrative, the prices are Anthropic’s first-party list rates, which Bedrock does not necessarily match, and the point is the ratios and the decisions.
Customer chat. About 1.5 million conversations a month, each roughly 3,000 tokens of context including account data and recent transactions, and 300 tokens of reply. That is payments data under the 2018 circular and personal data under DPDP, so it goes to the India profile: Sonnet 5 in-country, Guardrails in the same Region, invocation logging on. At $2 and $10 per million the bill is about $13,500 a month. Sonnet 5.5 offshore, on the same rate card with up to 30 percent fewer tokens per task and the 10 percent global discount, would be about $8,500. The residency premium is $5,000 a month and one generation of conversational quality; the legal review needed to send the data abroad would cost more than that in the first month. This is the easy case, and the one the announcement was built for.
Document extraction. Five million pages a month of KYC forms, loan files and statements, roughly 1,500 tokens in and 200 out per page. Personal data, so it stays in India, but bounded: fields, dates, amounts, a yes or no on completeness. On Haiku 4.5 in-country at $1 and $5 that is about $12,500 a month. On a permissively licensed small open model on a handful of the bank’s own GPUs it is the hardware and the engineers, which at this volume is less, and the model does not go away when a partner’s lifecycle table says so. For the decision steps, is this page a statement or a letter, does this signature match, a single-pass decision model with calibrated confidence runs in-country for a fraction of a cent per call, and Haiku handles the pages the small models cannot read. This is where open weights still win, and residency changed nothing.
Internal coding agent. Two thousand engineers, agentic sessions that re-read a 50,000-token context across 40 steps. The data is the bank’s source code: confidential, valuable, and neither personal nor payments data. Cache reads dominate. Using last week’s arithmetic, each run costs about $0.98 in context on Opus 5 at $0.50 cache reads and $0.39 on Opus 5.5 at $0.20. At ten runs per engineer per working day that is roughly $430,000 a month in-country against $170,000 offshore, for a model that is also a generation older on coding benchmarks. The right answer is the global profile, with two guards: a data-loss-prevention step that keeps customer records out of prompts, because engineers paste production logs, and a rule that routes any repository tagged as holding regulated data to the India profile automatically. Residency here is a per-repository setting, not a bank-wide policy.
Three workloads, three answers. Residency is now a routing choice, and the router has to know the data class of every request.
What to ask the vendor before signing
Because the product is a configuration value, the contract review is a list of questions about what the value guarantees. These are the ones we now put in writing.
- Which Regions can this profile route to, and can the list change without notice? AWS publishes the destination set for each geographic profile; Microsoft says it can add regions to a data zone without prior notice. Put the Region list in the contract.
- Where does flagged content go? Zero retention covers the normal path. Ask where abuse classification runs, whether a human can review a flagged prompt, in which country and under what retention.
- Who can read our prompts during a support case, and from where? Residency statements describe the automated path; support access is the manual one.
- When does the India profile retire, and when does its successor get one? Lifecycle dates on partner platforms are set by the partner. Ask for the schedule and for a commitment on time-to-India for new generations; for Opus 5.5 the answer today is not yet.
- What capacity is actually in the two Regions? Inference profiles do not support provisioned throughput. Get quota numbers for ap-south-1 and ap-south-2 and a statement of what happens when both are saturated.
- Who is the counterparty? The cloud provider, under which entity and governing law, with the lab’s terms as a marketplace licence. Make sure the audit-access and breach-notification clauses your regulator expects flow through to the model.
- What is the exit? Prompts, evaluations and guardrail configurations should be portable to another provider or to an open model on your own hardware. If the exit is a rewrite, the residency guarantee is only as good as the vendor’s next pricing decision.
Recommendations
- Treat residency as a data-class attribute, not an enterprise policy. Classify data so the router can decide per request whether a prompt may leave India, and default to in-country for anything you cannot classify.
- Buy the in-country frontier for regulated conversational and reasoning work now. Opus 5 and Sonnet 5 in Mumbai beat any model a bank could run itself with the same controls, and the compliance evidence is a log you already keep.
- Price the generation lag explicitly. Keep the global profile in the router for data that may leave, and report the cost of residency, in rate and in tokens per task, to the people who decide the data classes. On a cache-heavy agent it is two and a half times, not ten percent.
- Keep open weights for the bounded tier and for workloads where weights are the requirement. A small, permissively licensed model on your own GPUs, plus a decision model for the routing steps, is cheaper than any hosted option for extraction and classification, and it is the only option that survives a vendor’s lifecycle table.
- Write the five layers into the contract. Prompts, logs, control plane, weights, entity: for each, a sentence stating what is in India, what is not, and who decides. Residency is the first layer. Sovereignty is the paperwork for the other four.
Eight weeks ago I wrote that open weights were the only way to get frontier capability onto data that cannot leave. This week a bank in Mumbai can call a frontier model, one generation back, from two Regions in its own country, and prove it with a log. That is a real change, and a smaller one than the word sovereignty implies. The model is still theirs, the schedule is still theirs, and the capacity is where they put it. What India’s enterprises have gained is a choice they did not have in August, and the discipline it demands is the one every price cut and gated tier has demanded this year: know your data, route by class, and read the contract for what it does not say.