Built for the Real World · Essay · Sovereign AI

Claude answers from Mumbai: residency is a product now, and it is not sovereignty

On 5 October Anthropic made Claude Opus 5, Sonnet 5 and Haiku 4.5 available with in-country inference in India, routed between AWS’s Mumbai and Hyderabad Regions, seven weeks after AWS did the same for OpenAI’s GPT-5.6. Residency is now a feature the hosted labs sell. It costs a model generation and some price, and it covers prompts and logs, not weights, control plane or contract.

Abstract illustration for “Claude answers from Mumbai: residency is a product now, and it is not sovereignty”

On 5 October, Anthropic announced that Claude Opus 5, Claude Sonnet 5 and Claude Haiku 4.5 can now be called through an India endpoint on Amazon Bedrock, with every request processed inside the country. The plumbing had gone live on 29 September, when AWS published three India inference profiles, in.anthropic.claude-opus-5, in.anthropic.claude-sonnet-5 and in.anthropic.claude-haiku-4-5, which route each call between its Mumbai (ap-south-1) and Hyderabad (ap-south-2) Regions and nowhere else. Reliance and CRED tested the endpoint in a private preview. Kotak, Axis and IndusInd banks already use Claude, NPCI is building an agentic platform it calls AiNxt on it, TCS is rolling it out to 50,000 associates, Infosys has stood up an Anthropic centre of excellence, and Mahindra, Godrej, Swiggy, Zomato, Freshworks and Razorpay are customers. Irina Ghose, Anthropic’s managing director for India, said the company had “committed to bringing Claude inference to India” in August and was now delivering. Sandeep Dutta, who runs AWS in India and South Asia, supplied the commercial logic: in regulated industries, data residency can decide whether AI deployments “move from pilots into production”.

Seven weeks earlier AWS had done the same for OpenAI, adding India profiles for GPT-5.6 Terra and Luna on 17 August. Residency has become a product feature that the hosted labs sell, and they are shipping it to India within weeks of each other. In August I argued in a post on the trillion-parameter open-weight wave that open weights win on data that cannot leave, because the hosted option exists only where the provider has built a region. That was true in August and is less true in October. The short version of what follows: residency is now cheap to buy and costs you a model generation; and residency, even bought, is not sovereignty.

What shipped, and how the routing works

Bedrock has two kinds of cross-Region inference profile. A global profile lets AWS process a request in any commercial Region in the world and is priced about 10 percent below the alternative. A geographic profile restricts processing to Regions inside one geography, at standard pricing, and exists for compliance. The India profile is the geographic kind: a request sent to the Mumbai endpoint may be served in Mumbai or Hyderabad, depending on capacity, and the routing decision is AWS’s. The documentation is specific about the data. Traffic between the two Regions stays on the AWS network, encrypted in transit; customer data is “not stored in a destination Region”; CloudTrail records every request in the source Region, with an inferenceRegion field that says which of the two served it; Bedrock’s zero-data-retention default applies; Guardrails, invocation logging and the Converse and Messages APIs all work. Pricing is the standard Bedrock rate, charged to the source Region, with no routing surcharge.

The same release added Seoul profiles for Opus 5 and Sonnet 5 and a Singapore profile for Sonnet 5, so India was one of a batch. The India endpoint is not a sovereign cloud, a dedicated cluster or a special legal arrangement. It is a routing constraint on a shared fleet, delivered as a string in the model field, and its whole value is that it is boring.

Anthropic’s part is the model and the go-to-market: a Bengaluru office opened in February under Ghose, a partner network launched in March, a first partner certification programme in Bengaluru in August with a target of 5,000 people, and TCS as global premier partner. The OpenAI profiles, in.openai.gpt-5.6-terra and in.openai.gpt-5.6-luna, use the same two Regions and the same zero-retention model, with one visible difference: AWS’s OpenAI post says flagged content goes to offline abuse detection, and does not say where.

The generation lag

Here is the catch, and it is the most important thing in this post for anyone holding a budget. The India endpoint carries Opus 5, Sonnet 5 and Haiku 4.5. The global endpoint carries Opus 5.5, released on 22 September, Sonnet 5.5, released on 28 September, and Fable 5.1, released on 1 September. Opus 5.5 launched on Bedrock with US, EU, Australia, Japan and global profiles and no India profile; Sonnet 5.5 launched with the global profile only; Fable 5.1 has US and global profiles. So a bank that insists on in-country Claude inference is, on Anthropic’s own benchmark tables, one generation behind the bank next door that does not. I wrote last week about what that generation is worth: Sonnet 5.5 scores 70.6 on Terminal-Bench 4.0 against 66.4 for Opus 5.5, uses fewer tokens for the same work and generates output more than 30 percent faster than Sonnet 5; Opus 5.5 is 20 percent cheaper than Opus 5 on list and 60 percent cheaper on cache reads, $0.20 against $0.50 per million. Those are first-party list prices, not Bedrock’s, but the shape of the gap carries over.

So the price of residency is concrete: roughly 10 percent on rate against the global profile, plus the per-task efficiency, the cache cut and the benchmark points of the generation you cannot have. For an agentic workload whose bill is mostly cache reads, in-country Opus 5 costs about two and a half times what offshore Opus 5.5 costs for the same context. For a chat workload on Sonnet, where the list price did not change, the cost is the up to 30 percent fewer tokens per task that you do not get.

This is not an Anthropic quirk; it is how every hosted lab ships. OpenAI’s India profiles carry GPT-5.6 Terra and Luna; GPT-6 Sol and Luna reached Bedrock on 22 September, and the announcement lists no India profile. Microsoft’s documentation for Azure OpenAI is unusually candid: new models arrive as Global deployments first, then Data Zone, then geography-based, and the geography-based types “have no guaranteed availability date”. Claude on Vertex AI is the starkest case. Anthropic’s own documentation says regional endpoints support Sonnet 4.6 and earlier, that newer models use the global endpoint or the US and EU multi-region endpoints, and that regional and multi-region endpoints carry a 10 percent premium; there is no India multi-region endpoint. Capacity for a new model is scarce at launch, the labs put it where demand is densest, and a two-Region geography on the far side of the world gets it when it gets it. Table 1 sets the offers side by side.

OfferWhere inference runsStorage at rest and logsModel generation on 6 Oct 2026How residency is enforcedPricing notes
Claude on Amazon Bedrock, India profile
AWS · 29 Sep; Anthropic · 5 Oct
Mumbai (ap-south-1) and Hyderabad (ap-south-2) only; AWS chooses between them per requestNot stored in the destination Region; zero retention by default; CloudTrail and CloudWatch write in the source RegionOpus 5, Sonnet 5, Haiku 4.5. Global profile: Opus 5.5, Sonnet 5.5, Fable 5.1in.anthropic. inference profile ID, plus an organisation policy allowing only the two RegionsStandard rate, billed to the source Region; global profile about 10% cheaper
OpenAI on Amazon Bedrock, India profile
AWS · 27 Aug
The same two RegionsZero retention; flagged content goes to offline abuse detection, location not statedGPT-5.6 Terra and Luna. GPT-6 Sol and Luna reached Bedrock on 22 Sep with no India profile in the announcementin.openai. inference profile IDStandard rate per the Bedrock pricing page
OpenAI direct (API, ChatGPT Enterprise)
OpenAI · May 2025
No in-country processing commitment; residency is for storageData at rest in India for eligible API customers and new Enterprise and Edu workspacesNewest: GPT-6 Sol, Luna and AstraWorkspace or project residency settingList prices. The Tata agreement of 18 Feb 2026, 100 MW scaling to 1 GW, points to in-country capacity with no date
Azure OpenAI, Standard (regional) deployment
Microsoft docs · updated Aug 2026
Within the chosen Azure geography, and may move between regions inside it for operations; Data Zone types process within the US, EU or APAC zoneAt rest in the resource’s geography for every deployment typeArrives Global first, then Data Zone, then geography, which has no guaranteed availability dateDeployment SKU (Standard rather than GlobalStandard); Azure Policy can block the global SKUsGlobal Standard has the lowest price; Microsoft can add regions to a data zone without notice
Claude on Vertex AI, regional endpoint
Anthropic docs · Oct 2026
The named region; multi-region endpoints exist for the US and EU onlyGoverned by Google Cloud; request-response logging is optionalRegional endpoints support Sonnet 4.6 and earlier; Opus 5.5, Sonnet 5.5 and Fable 5.1 use the global or US and EU endpointsRegional endpoint URL10% premium on regional and multi-region endpoints over global
Open weights, self-hosted
see the 14 Aug post
Your racks, or a cloud tenancy you chooseYours, with your logs and your keysQwen3.8-Max (no licence at release), Kimi K3, Inkling, Nemotron 3.5 Lightning; a few points behind the hosted frontierPhysics, and your egress policyHardware plus two to three platform engineers; a six-figure monthly line at trillion scale, a single GPU for Nemotron-class models
Table 1. The residency offers an Indian enterprise can buy on 6 October 2026, from the providers’ own documentation and announcements. The model-generation column is the one to read first: every hosted offer that keeps processing in a chosen place is at least one generation behind the same provider’s global endpoint.

What residency covers, and what it does not

Residency, as sold, covers one thing: where the compute that runs inference sits, and therefore where your prompts and outputs exist in plaintext for the milliseconds that matter. On Bedrock’s India profile that commitment is clean, and so is the logging, because CloudTrail and CloudWatch write in the Region you called. Beyond that ring the picture changes. Figure 1 lays it out as five layers from the inside out.

Prompts and outputs. Covered. Processed in Mumbai or Hyderabad, not stored in the destination Region, zero retention by default.

Logs and telemetry. Partly covered. Your invocation logs and audit trail stay in the source Region. What is not stated, for either lab, is where safety and abuse classification runs when a request is flagged, who reviews it and under which jurisdiction. For most workloads that is a footnote; for a bank’s customer transcripts it is a question to ask in writing.

Control plane. Not covered. Which Regions the profile may route to, the quota in each, the model catalogue, the Guardrails service, the release train that decides when Opus 5.5 gets an India profile: all of it belongs to AWS and Anthropic. Inference profiles do not support Provisioned Throughput, so the capacity you get in India is the capacity AWS chose to put there.

Weights. Not covered, and never will be on a hosted offer. The model file is Anthropic’s, on AWS hardware, in AWS accounts; you cannot inspect it, copy it, fine-tune it in place or pin it against a retirement. This is the layer open weights give you and nothing else does, and the layer most of India’s sovereign-AI conversation is actually about.

Legal entity and terms. Not covered. Your contract is with the cloud provider, and the model reaches you through its marketplace terms; the lab is not your counterparty. Lifecycle is the case in point. Anthropic’s platform documentation says lifecycle dates on partner-operated platforms are “set by the partner and can differ from the Claude API schedule”. The India profile you adopt today has a retirement date you have not been told, set by a company you may have no direct relationship with, and its successor has no India profile yet.

What in-country inference covers, and what it does notFive nested layers from the inside out: prompts and outputs, logs and telemetry, control plane, weights, and the legal entity and terms. The India endpoint on Amazon Bedrock covers the innermost layer fully and the logging layer partly; the control plane, the weights and the contracting entity remain with the cloud provider and the lab. FIVE LAYERS, INSIDE OUTCLAUDE ON BEDROCK, INDIA PROFILE, 6 OCT 2026 Legal entity and terms Weights Control plane Logs and telemetry Prompts and outputs processed in ap-south-1 or ap-south-2 not stored in the destination Region Prompts and outputs Covered. Inference in Mumbai or Hyderabad, encrypted in transit on the AWS network; zero data retention by default. Logs and telemetry Partly. CloudTrail and CloudWatch log in the source Region. Where flagged content is reviewed, and by whom, is unstated. Control plane Not covered. Routing, quotas, catalogue and the release train are AWS’s and Anthropic’s; no provisioned throughput. Weights Not covered. Anthropic’s file on AWS hardware; you cannot inspect, copy, fine-tune or pin it against retirement. Legal entity and terms Not covered. Contract with the cloud provider; lab terms as a marketplace licence; lifecycle dates set by the partner.
Figure 1. Residency is the innermost layer. Slate: covered by the India profile as documented by AWS. Clay: remains with the cloud provider and the lab. Sovereignty, in the sense a regulator or a board means it, is the whole stack, and a hosted offer can only ever sell the inside two rings.

None of this is an argument against buying residency; it is an argument for naming what you bought. In August I listed five layers of sovereignty, weights, licence, compute, data and controls, and said open weights give you one. Hosted residency gives you a different one, the data layer, with the lab’s controls thrown in. They still do not add up to five.

The regulatory map that makes this matter in India

Indian regulation is layered, and the layers point in different directions.

DPDP. The Digital Personal Data Protection Act of 2023 is, on its face, permissive about transfers. Section 16 lets the central government restrict transfers to countries it names, a negative list, and section 16(2) preserves any other law that imposes a higher degree of protection. The Rules notified in November 2025 keep that structure and add the thing that bites: Rule 12(4) lets the government specify categories of personal data that Significant Data Fiduciaries must keep in India, along with the related traffic data, and the largest banks and platforms should expect that designation. The substantive obligations, including 72-hour breach notification to the Data Protection Board and penalties of up to ₹250 crore, take effect in May 2027. DPDP does not mandate localisation of inference; it creates a switch under which localisation of named categories can be turned on, and it leaves every sectoral rule in force.

RBI and payments data. The sectoral rule that matters most is older and sharper. The RBI’s circular of 6 April 2018 required that end-to-end transaction data for payment systems be stored only in India, with compliance by 15 October that year. Its FAQ of 26 June 2019 allowed processing abroad provided the data is deleted from the foreign system and brought back to India “not later than the one business day or 24 hours” after processing, whichever is earlier, and extended the obligation to banks, payment system operators and the third-party vendors they engage. Against a hosted model offshore, every prompt containing a UPI transaction is payments data processed abroad, and you owe the regulator evidence that it was deleted within 24 hours from an inference stack you do not operate. In-country inference does not change the rule; it removes the proof burden, which in practice is the whole difficulty.

RBI on AI itself. On 13 August 2025 the RBI published the report of its FREE-AI committee, chaired by Pushpak Bhattacharyya of IIT Bombay: seven guiding sutras and 26 recommendations across six pillars, among them financial-sector data infrastructure, indigenous models for financial services, an innovation sandbox, a standard AI incident report and clearer liability. It is advisory, but it signals where supervision is going: a bank will be asked not only where its data is but which model made the decision and who answers for it.

Securities and government. SEBI’s cloud framework of 6 March 2023 requires a regulated entity’s data to be stored and processed in data centres of cloud providers empanelled by MeitY, with the entity retaining ownership of its data, logs and keys. For the public sector, which AWS and Anthropic named as a target segment, processing location is the first gate in any procurement.

Nothing in Indian law says a large language model must run in India. Several things say particular data must stay in India, or come back within a day, or be kept under conditions the government can tighten by notification. Residency as a product is valuable because it collapses a dozen attestation problems into one configuration value a regulator can read in a CloudTrail log.

Three options for a regulated Indian enterprise

The choice for any given workload comes down to three postures, compared on cost, capability, control and compliance.

Open weights on your own GPUs. Cost: hardware, two to three platform engineers and, for a trillion-scale model such as Kimi K3, a six-figure monthly infrastructure line before a token is served; for a small model such as Nemotron 3.5 Lightning, a single GPU. Capability: a few benchmark points behind the hosted frontier at the top, level with it on bounded tasks. Control: complete, including the whole job of containment. Compliance: trivially satisfied on location, but the licence has to clear, and Qwen3.8-Max shipped without one.

In-country hosted frontier. Cost: standard Bedrock rates, roughly 10 percent above the global profile, on a model a generation back. Capability: Opus 5, Sonnet 5 and Haiku 4.5, or GPT-5.6 Terra and Luna. Control: the lab’s controls as a service, none of the weights, a lifecycle you do not set. Compliance: processing location, retention and logging are clean and evidenced; the classifier and support questions need a written answer.

Offshore hosted frontier. Cost: the cheapest per task, with the newest cache pricing. Capability: Opus 5.5, Sonnet 5.5, Fable 5.1, GPT-6 Sol and Luna, and whatever ships next week, including the gated tiers no in-country endpoint will see for a long time. Control: as above. Compliance: only for data classes that may leave, a determination your data classification must be able to make per request, not per project.

Notice what the in-country option did to the open-weights case. In August the argument for self-hosting a 2.4-trillion-parameter model was that for regulated data there was no alternative. Now there is one, for most regulated data, at a fraction of the operational effort. Open weights retreat to where they were always strongest: the bounded, high-volume tier where a 3.6-billion-active-parameter model on one GPU beats any API on price, and the workloads where the weights themselves are the requirement, because you need to fine-tune, pin a version for seven years or run air-gapped. Smaller territory than eight weeks ago, and still real.

A worked example: one bank, three workloads

Take a mid-sized private bank with three things on the AI roadmap. The token counts are ours and illustrative, the prices are Anthropic’s first-party list rates, which Bedrock does not necessarily match, and the point is the ratios and the decisions.

Customer chat. About 1.5 million conversations a month, each roughly 3,000 tokens of context including account data and recent transactions, and 300 tokens of reply. That is payments data under the 2018 circular and personal data under DPDP, so it goes to the India profile: Sonnet 5 in-country, Guardrails in the same Region, invocation logging on. At $2 and $10 per million the bill is about $13,500 a month. Sonnet 5.5 offshore, on the same rate card with up to 30 percent fewer tokens per task and the 10 percent global discount, would be about $8,500. The residency premium is $5,000 a month and one generation of conversational quality; the legal review needed to send the data abroad would cost more than that in the first month. This is the easy case, and the one the announcement was built for.

Document extraction. Five million pages a month of KYC forms, loan files and statements, roughly 1,500 tokens in and 200 out per page. Personal data, so it stays in India, but bounded: fields, dates, amounts, a yes or no on completeness. On Haiku 4.5 in-country at $1 and $5 that is about $12,500 a month. On a permissively licensed small open model on a handful of the bank’s own GPUs it is the hardware and the engineers, which at this volume is less, and the model does not go away when a partner’s lifecycle table says so. For the decision steps, is this page a statement or a letter, does this signature match, a single-pass decision model with calibrated confidence runs in-country for a fraction of a cent per call, and Haiku handles the pages the small models cannot read. This is where open weights still win, and residency changed nothing.

Internal coding agent. Two thousand engineers, agentic sessions that re-read a 50,000-token context across 40 steps. The data is the bank’s source code: confidential, valuable, and neither personal nor payments data. Cache reads dominate. Using last week’s arithmetic, each run costs about $0.98 in context on Opus 5 at $0.50 cache reads and $0.39 on Opus 5.5 at $0.20. At ten runs per engineer per working day that is roughly $430,000 a month in-country against $170,000 offshore, for a model that is also a generation older on coding benchmarks. The right answer is the global profile, with two guards: a data-loss-prevention step that keeps customer records out of prompts, because engineers paste production logs, and a rule that routes any repository tagged as holding regulated data to the India profile automatically. Residency here is a per-repository setting, not a bank-wide policy.

Three workloads, three answers. Residency is now a routing choice, and the router has to know the data class of every request.

What to ask the vendor before signing

Because the product is a configuration value, the contract review is a list of questions about what the value guarantees. These are the ones we now put in writing.

  1. Which Regions can this profile route to, and can the list change without notice? AWS publishes the destination set for each geographic profile; Microsoft says it can add regions to a data zone without prior notice. Put the Region list in the contract.
  2. Where does flagged content go? Zero retention covers the normal path. Ask where abuse classification runs, whether a human can review a flagged prompt, in which country and under what retention.
  3. Who can read our prompts during a support case, and from where? Residency statements describe the automated path; support access is the manual one.
  4. When does the India profile retire, and when does its successor get one? Lifecycle dates on partner platforms are set by the partner. Ask for the schedule and for a commitment on time-to-India for new generations; for Opus 5.5 the answer today is not yet.
  5. What capacity is actually in the two Regions? Inference profiles do not support provisioned throughput. Get quota numbers for ap-south-1 and ap-south-2 and a statement of what happens when both are saturated.
  6. Who is the counterparty? The cloud provider, under which entity and governing law, with the lab’s terms as a marketplace licence. Make sure the audit-access and breach-notification clauses your regulator expects flow through to the model.
  7. What is the exit? Prompts, evaluations and guardrail configurations should be portable to another provider or to an open model on your own hardware. If the exit is a rewrite, the residency guarantee is only as good as the vendor’s next pricing decision.

Recommendations

  1. Treat residency as a data-class attribute, not an enterprise policy. Classify data so the router can decide per request whether a prompt may leave India, and default to in-country for anything you cannot classify.
  2. Buy the in-country frontier for regulated conversational and reasoning work now. Opus 5 and Sonnet 5 in Mumbai beat any model a bank could run itself with the same controls, and the compliance evidence is a log you already keep.
  3. Price the generation lag explicitly. Keep the global profile in the router for data that may leave, and report the cost of residency, in rate and in tokens per task, to the people who decide the data classes. On a cache-heavy agent it is two and a half times, not ten percent.
  4. Keep open weights for the bounded tier and for workloads where weights are the requirement. A small, permissively licensed model on your own GPUs, plus a decision model for the routing steps, is cheaper than any hosted option for extraction and classification, and it is the only option that survives a vendor’s lifecycle table.
  5. Write the five layers into the contract. Prompts, logs, control plane, weights, entity: for each, a sentence stating what is in India, what is not, and who decides. Residency is the first layer. Sovereignty is the paperwork for the other four.

Eight weeks ago I wrote that open weights were the only way to get frontier capability onto data that cannot leave. This week a bank in Mumbai can call a frontier model, one generation back, from two Regions in its own country, and prove it with a log. That is a real change, and a smaller one than the word sovereignty implies. The model is still theirs, the schedule is still theirs, and the capacity is where they put it. What India’s enterprises have gained is a choice they did not have in August, and the discipline it demands is the one every price cut and gated tier has demanded this year: know your data, route by class, and read the contract for what it does not say.

Sources

  1. DataQuest India. Anthropic brings in-country Claude inference to India via Amazon Bedrock. 5 October 2026.
  2. CRN Asia. Anthropic takes Claude inference live in India, opens path to regulated AI deployments. October 2026.
  3. AWS Machine Learning Blog. Amazon Bedrock expands Claude model availability to in-country inferencing in India. 29 September 2026.
  4. AWS Machine Learning Blog. Introducing OpenAI models on Amazon Bedrock for in-country inferencing in India. 20 August 2026.
  5. TechCrunch. OpenAI taps Tata for 100MW AI data center capacity in India, eyes 1GW. 18 February 2026.
  6. Microsoft Learn. Understanding deployment types in Microsoft Foundry Models. Updated August 2026.
  7. Reserve Bank of India. FAQs: Storage of Payment System Data. 26 June 2019.
  8. KPMG India. RBI’s FREE-AI committee report in the financial sector. September 2025.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn