Built for the Real World · Essay · Sovereign AI

LLM data residency in India: what runs in-country in October 2026, and where it breaks

Amazon Bedrock now runs Claude, GPT-5.6 and Kimi K3 inside India, Azure OpenAI processes a handful of older GPT models in South India, and Vertex AI commits to Mumbai for two Gemini Flash models. This is the October 2026 map of LLM data residency in India, provider by provider, and of the hops where residency claims break: gateways, logs and abuse monitoring.

Abstract illustration for “LLM data residency in India: what runs in-country in October 2026, and where it breaks”

On 6 October, AWS added Moonshot’s Kimi K3 to the short list of models that Amazon Bedrock will run only inside India. Requests sent to the new in.moonshotai.kimi-k3 inference profile are processed in Mumbai or Hyderabad and nowhere else. Kimi K3 is the third model family with an India profile, after OpenAI’s GPT-5.6 Terra and Luna in August and Anthropic’s Claude Opus 5, Sonnet 5 and Haiku 4.5 on 29 September. On 7 October, Google updated the data residency page that states where Gemini on Vertex AI processes data. Its India column lists Gemini 3.5 Flash and Gemini 2.5 Flash, and none of the three Flash releases that followed them.

So the question has changed. LLM data residency in India is no longer a matter of whether a hosted model can run in the country: on three hyperscalers and several Indian GPU clouds, it can. The questions now are which models, of which generation, and which hops of a request actually stay. This post is the map, built from the providers’ own documentation as it stood this week. The Claude detail, and what its generation lag costs, is in yesterday’s post. The argument here is broader: residency claims rarely break at the model. They break at the hops around it.

What LLM data residency in India actually covers

Data at rest. Uploaded files, fine-tuned weights, stored conversation state, batch inputs waiting to run. All three hyperscalers keep this in the region or geography you chose, whatever deployment type you pick. It is the easy part, and what most “stored in India” announcements describe.

Inference processing. Where the accelerator that runs the model sits, and so where your prompt and its answer exist in plaintext. This is what inference profiles, deployment types and endpoints change, and where the products differ. A global option on any of the three clouds can process a request wherever the provider chooses. Google is the bluntest: its endpoints “don’t guarantee data residency or in-region ML processing”, and a separate per-model table lists the commitments it does make.

Logs and safety copies. Your own logs, audit trails and traces, which you configure. And the provider’s abuse-monitoring copies, which you do not, and which vary by model. Amazon Bedrock stores no inputs or outputs by default, but keeps classifier-flagged traffic for OpenAI’s GPT-5.6 and GPT-6 models for up to 30 days, and all traffic for Claude Fable 5 and 5.1 for up to 30 days, in the Region that served the request. Azure stores flagged samples for possible human review in your resource’s geography. Google may log flagged prompts for up to 90 days in your project’s region, unless you are on a Google Cloud Master Agreement, which is exempt by default.

Providers sell the second item. Regulators mostly ask about the first and the third. The gap between them is where residency claims fail.

What Indian rules actually require

The common version of the obligation, that Indian law requires AI to run in India, is wrong. No Indian rule says a language model must run in the country. Several say that particular data must be stored, processed or logged here, and they differ by sector.

The DPDP Act and Rules. The Digital Personal Data Protection Act of 2023 permits transfers abroad unless the government restricts a country by notification, and preserves any sectoral law that sets a higher bar. The Rules notified on 13 November 2025 keep that shape. Rule 15 allows transfers subject to requirements the government may set, by order, about making data available to a foreign state. Rule 13(4) is the one to plan for: it lets the government specify, on a committee’s recommendation, personal data that a Significant Data Fiduciary must keep in India, together with the traffic data about its flow. Both come into force in May 2027.

RBI on payment data. The Reserve Bank’s circular of 6 April 2018 requires the entire data relating to payment systems to be stored only in India. Its 2019 FAQ allows processing abroad if the data is deleted there and brought back within one business day or 24 hours, whichever is earlier, and extends the rule to the service providers banks use. A prompt that carries a customer’s transactions is payment data.

SEBI on cloud. SEBI’s cloud framework of 6 March 2023 is the strictest of the set. For exchanges, depositories, brokers, asset managers and its other regulated entities, data in the cloud, logs included, should reside and be processed “within the legal boundaries of India”, in data centres of MeitY-empanelled providers with a valid STQC audit. A hosted model is a cloud service, so sending such data to a model abroad is hard to square with that text.

Insurance. IRDAI’s 2023 guidelines require intermediaries, though not insurers, to keep ICT logs and critical data in India; insurers must keep primary data here under record-keeping rules.

Logs, for everyone. CERT-In’s directions of 28 April 2022 require service providers, data centres, body corporates and government organisations to keep logs of all ICT systems for a rolling 180 days, within the Indian jurisdiction. An LLM gateway is an ICT system.

So the binding constraints attach to data classes, not models: payment data, securities-market data, insurance records and logs, plus a DPDP switch for named categories from 2027. Residency has to be decided per request, not per vendor.

From Mumbai, Bedrock’s APAC profile can route to Tokyo, Seoul, Osaka, Singapore and Sydney. APAC is not India

Amazon Bedrock in Mumbai and Hyderabad: three ways to call a model

AWS has two Indian Regions, Mumbai (ap-south-1) and Hyderabad (ap-south-2), and Bedrock offers three ways to call a model from them. In-Region inference processes the request in the Region you call. A geographic profile routes it among Regions in one geography: the US, the EU, APAC, Japan, Australia or India. A global profile sends it to any commercial Region, about 10 percent cheaper for some models. The model ID says which: a bare ID for in-Region, a prefix such as in. or apac. for a geography, global. for the rest.

This week the India profile carries six models: Claude Opus 5, Sonnet 5 and Haiku 4.5; GPT-5.6 Terra and Luna; and Kimi K3. Each routes only between Mumbai and Hyderabad, and CloudTrail records, written in the Region you called, note which of the two served each request. The newer models are missing. Claude Opus 5.5, Sonnet 5.5 and Fable 5.1, and OpenAI’s GPT-6 Sol, Luna and Astra, reach the Indian Regions only through the global profile.

In-Region inference in Mumbai is a different catalogue, and the strictest residency Bedrock sells. It is almost entirely open-weight: gpt-oss-120b, Qwen3 235B and Qwen3 Coder 480B, DeepSeek V3.2, Mistral Large 3, Kimi K2.5, GLM 5, MiniMax M2.5, Gemma 3 and NVIDIA’s Nemotron Nano family, most of which also run batch jobs in Mumbai alone. Hyderabad’s in-Region list is a single embedding model.

Then there is the trap. Amazon’s Nova Pro, Lite and Micro are offered from Mumbai through a geographic profile, but it is the APAC one, and from Mumbai it routes to Tokyo, Seoul, Osaka, Singapore and Sydney as well. It is priced like any geographic profile and satisfies none of the Indian rules above. AWS’s advice is to call GetInferenceProfile from every source Region and check every destination, rather than rely on “the geographic prefix in an inference profile ID alone”. Enforce it with a service control policy that denies Bedrock inference outside the two Indian Regions: a cross-Region request fails if any destination is blocked, so a pasted global or APAC ID returns an error, not an offshore answer.

Azure OpenAI in India: South India is the region, and the data zone is not India

Microsoft has four cloud regions in India: Central India in Pune, South India in Chennai, West India in Mumbai, and India South Central in Hyderabad, live since 6 August with three availability zones. For Azure OpenAI only one matters today: Microsoft’s availability tables for Foundry Models list South India and no other Indian region.

Azure’s residency choice is the deployment type. Global types may process a prompt in any Azure region. Data Zone types process within the United States, the European Union or Asia Pacific, and Microsoft defines the last as any Asia Pacific nation. Standard and Regional Provisioned deployments process in the deployment’s region, moving within the geography only for operational reasons. Only those two keep inference in India.

What they carry is the question. Standard in South India offers gpt-4.1-mini and gpt-4o, plus embeddings and Whisper. Regional Provisioned, where you reserve throughput units, reaches gpt-5, gpt-5.1 and gpt-5.4-mini. GPT-5.6 and the GPT-6 models are offered from South India only as Global or Asia Pacific Data Zone deployments. Batch has no regional type at all, so a batch deployment in South India is Global Batch and can run anywhere. Microsoft is candid that geography-based types get new models last.

On abuse monitoring, flagged prompts and completions are stored in your resource’s geography and may be reviewed by Microsoft employees. For the European Economic Area the documentation says those reviewers are in the EEA; it says nothing equivalent for India. Eligible customers can apply for modified abuse monitoring, which removes the storage and the human review.

Vertex AI and Gemini in Mumbai: two Flash models and a commitment table

Google Cloud has two Indian regions, Mumbai (asia-south1) and Delhi (asia-south2). For generative AI, Delhi appears in neither the endpoint list nor the residency table, so the map is Mumbai alone. Google, which now documents Vertex AI under its Gemini Enterprise Agent Platform, separates endpoints from commitments: the data residency page states, model by model, where ML processing is guaranteed.

The India column of that table has three entries: Gemini 3.5 Flash, the 128,000-token configuration of Gemini 2.5 Flash, and text-embedding-005. Gemini 3.6, 3.7 and 3.8 Flash carry commitments only for the US and EU multi-regions. No Gemini Pro model has an India commitment. Nor does any partner or open model: Claude on Google Cloud has commitments in Singapore, Belgium or Taiwan for some versions, and none in India.

Two Google features create copies. Grounding with Google Search stores queries derived from prompts for up to three days; if you use it, that storage cannot be switched off, and the documentation does not say where the logs are kept. And Gemini caches inputs and outputs in memory for up to 24 hours by default, within the selected location, which can be turned off per project.

OptionIndian RegionsModels processed in India, 7 Oct 2026What else stays in IndiaCaveats
Amazon Bedrock, India profile (in.)
AWS · Aug to Oct 2026
Mumbai (ap-south-1) and Hyderabad (ap-south-2); AWS picks per requestClaude Opus 5, Sonnet 5, Haiku 4.5; GPT-5.6 Terra and Luna; Kimi K3 (from 6 Oct)CloudTrail and invocation logs in the source Region; any retained safety copy in the Region that served the callOpus 5.5, Sonnet 5.5, Fable 5.1 and the GPT-6 models are global-only from India; GPT-5.6 flagged traffic kept up to 30 days; no Provisioned Throughput on profiles
Amazon Bedrock, in-Region
AWS · Mumbai
Mumbai only for text modelsgpt-oss-120b, Qwen3 235B, Qwen3 Coder 480B, DeepSeek V3.2, Mistral Large 3, Kimi K2.5, GLM 5, MiniMax M2.5Everything; the request never leaves the RegionOpen-weight models, a tier below the hosted frontier; most support single-Region batch
Amazon Bedrock, APAC profile (apac.)
AWS · not residency
From Mumbai: Tokyo, Seoul, Osaka, Singapore, Sydney and MumbaiNone guaranteed. Nova Pro, Lite and Micro use this routeYour logs onlySame price as any geographic profile; satisfies no Indian residency rule
Azure OpenAI, Standard and Regional Provisioned
Microsoft · South India (Chennai)
South India; may move within the geography for operationsStandard: gpt-4.1-mini, gpt-4o. Reserved capacity: up to gpt-5.1 and gpt-5.4-miniData at rest; abuse-monitoring store in the resource’s geographyData Zone means Asia Pacific; batch is Global only; GPT-5.6 and GPT-6 not offered in-region; Hyderabad region not yet listed
Gemini on Vertex AI, regional endpoint
Google · Mumbai (asia-south1)
Mumbai; Delhi (asia-south2) not listedGemini 3.5 Flash; Gemini 2.5 Flash (128k context); text-embedding-005Data at rest; flagged-prompt logs up to 90 days in-region; 24-hour in-memory cacheGemini 3.6 to 3.8 Flash, all Pro models, Claude and open models have no India commitment; Search grounding logs kept 3 days, location unstated
Lab APIs direct
Anthropic, OpenAI
None for processingNoneOpenAI: storage at rest in India for eligible customers since May 2025Anthropic’s inference_geo offers “us” or “global” only
Indian GPU clouds and IndiaAI compute
Yotta, E2E Networks and others
Providers’ Indian data centresOpen-weight catalogues (Llama, DeepSeek, Qwen and others) or your own weightsEverything, per the provider’s termsHosting location and audits are per facility, so get both in writing; subsidised IndiaAI capacity targets start-ups and academia
Open weights, self-hosted
your tenancy or racks
Any Indian region, or on premisesAny model whose licence clearsEvery hop, if you build it that wayYou own capacity, patching, safety filtering and licence review
Table 1. Where an Indian enterprise can run a hosted or open model with processing in India, from AWS, Microsoft, Google and provider documentation as of 7 October 2026. Read the third column against the newest model each provider sells: every hosted in-country option trails the same provider’s global catalogue, and the APAC routes are not Indian residency at all.

Sovereign AI cloud in India: GPU clouds, IndiaAI compute and your own weights

The fourth option is to take the model to a GPU you choose.

Indian GPU clouds. Yotta’s Shakti Studio, promoted with a free-credit programme in August, offers serverless access to more than 25 open models through OpenAI-compatible endpoints; Yotta says it is hosted entirely in its own Indian data centres and billed in rupees. E2E Networks’ TIR platform offers OpenAI-compatible endpoints for open models such as Llama and DeepSeek, billed per request, plus dedicated endpoints for your own weights; its GenAI documentation does not say where the hosted endpoints run, which is the first question to ask. For a SEBI-regulated buyer, the second is whether the specific facility is MeitY-empanelled.

The IndiaAI compute portal. The government said in March that more than 38,000 GPUs had been onboarded through the IndiaAI Mission’s portal, offered to Indian start-ups and academia at subsidised rates. Figures given in the Lok Sabha in August put the empanelled providers at 15; Yotta, E2E Networks and NxtGen were the first three integrated with the portal, by April 2025, and the hardware spans NVIDIA H100, H200 and A100, AMD MI300X, Intel Gaudi 2 and AWS Trainium. It is national capacity for building models, not an enterprise inference service, though the same providers sell to enterprises directly.

Your own weights. An open-weight model on GPUs in an Indian cloud region, or in your own racks, keeps every hop in the country by construction, if you build the gateway, logs and monitoring the same way. Bedrock’s in-Region catalogue is the managed version. The self-managed version is the one I described in August’s post on trillion-scale open weights: you own capacity, patching, safety filtering and the licence review, and a trillion-parameter model needs a multi-node cluster before it serves a token. Kimi K3 now comes both ways: as an India profile on Bedrock and as a download.

Where residency claims break

When the platform team I run looks at a residency claim, the first step is to draw the path of one request and mark each hop. Figure 1 is that drawing, generalised. The option you buy decides one hop, inference, and partly decides a second, the provider’s safety copy. Every other hop is yours, and that is where the breaks are.

The gateway. If the router in front of your models runs in a Singapore region, or is a hosted service with its backend abroad, every prompt leaves India before it reaches the Mumbai endpoint.

Traces. Observability tools record full prompts and completions by design; one that ships spans to a hosted backend abroad keeps an offshore copy of every conversation, on someone else’s retention schedule.

The safety copy. Model-specific, as above. The same Bedrock India profile is zero-retention for Claude Opus 5 and keeps flagged GPT-5.6 traffic for up to 30 days, in India in both cases. On a global profile, any retained copy sits wherever the request was processed.

Add-on features. Web grounding, hosted retrieval stores, stored conversation state and agent memory create data the inference commitment does not cover. Search grounding is the clearest example; Azure, which keeps stored completions and Responses API state at rest in the resource’s geography, is the good case.

Fallbacks. A router that fails over to a global endpoint when the India quota runs out turns a residency guarantee into a preference. Bedrock’s inference profiles do not support Provisioned Throughput, so plan the India quota explicitly.

Which hops of an LLM request stay in India, by residency optionA grid of four residency options against five hops of a request: gateway, inference, provider logs, provider safety copy and add-on features. Global options and regional APAC zones let inference leave India; in-country options keep inference, logs and safety copies in India; self-hosted open weights keep every hop in India. The gateway and add-ons are decided by the customer’s own configuration in every hosted option. FOUR OPTIONS, FIVE HOPS OF ONE REQUESTAS DOCUMENTED, 7 OCT 2026 GatewayInferenceProvider logsSafety copyAdd-ons router and keysmodel computeaudit, invocationabuse monitoringgrounding, tracing Global option Regional zone In-country option Open weights, self-hosted Bedrock global., Azure Globaland Global Batch, Vertex global Bedrock apac. profiles, AzureAsia Pacific Data Zone Bedrock in. and in-Region, AzureSouth India, Vertex asia-south1 Indian GPU cloud, an Indiancloud region, or your racks Your choiceAnywhereIndiaVariesCheck each Your choiceAcross APACIndiaVariesCheck each Your choiceIndiaIndiaIndiaCheck each IndiaIndiaIndiaNone, or yoursYour choice Stays in India, as documented Can leave India Decided by your own configuration
Figure 1. Where one request’s data goes, by residency option, from AWS, Microsoft and Google documentation as of 7 October 2026. The option you buy moves the inference column; on global and regional options, where the safety copy sits depends on the provider and the model. The tinted cells are set by your own gateway, tracing and feature choices, and that is where most residency claims break.

A worked example: one bank, three workloads

Take an illustrative large private bank with a broking subsidiary and three workloads on its roadmap.

A customer-facing assistant. It answers account and card questions in the app, which puts names, account numbers and recent transactions into the prompt: personal data under DPDP and payment data under the RBI circular. Evidencing deletion from a provider’s inference fleet within a day is impractical, so every hop stays in India. The candidates are Bedrock’s India profile, with Claude Sonnet 5 or Opus 5, GPT-5.6 Terra or Kimi K3; Azure in South India, with gpt-4.1-mini on Standard or gpt-5.1 on reserved capacity; and Gemini 3.5 Flash in Mumbai. The bank’s own evaluation chooses among them.

The residency work is the same whichever wins: a gateway in Mumbai, gateway and invocation logs kept in India for at least 180 days, a policy that denies every non-Indian route, and no web grounding. Then the safety-copy terms, where the clouds differ. On Bedrock, Claude Opus 5 and Sonnet 5 are zero-retention. On Azure, flagged samples are kept for human review unless the bank is approved for modified abuse monitoring. On Google, flagged prompts may be logged unless the bank has a master agreement or an exception. Choose the option whose safety copy the bank can describe to its regulator in one sentence.

An internal coding assistant. The data is source code and questions about it: confidential, but usually neither personal nor payment data, so neither DPDP nor the RBI circular reaches it. This is where the newest generation is worth most, and the bank can use global options, which are also the cheapest. What it should not use is the middle: an APAC profile or an Asia Pacific data zone gives neither Indian residency nor the lowest price.

Three things still stay in India. The gateway and its logs, under CERT-In. A redaction step for the production logs and customer records engineers paste into prompts, which become payment data the moment they contain a transaction. And the coding tool’s own backend: many assistants send code to the vendor’s servers before any model sees it, so choose one that can be pointed at the bank’s gateway. The broking subsidiary is the exception. SEBI’s framework covers any data pertaining to the regulated entity in the cloud, and its compliance team may read that to include code and logs. If it does, the subsidiary’s engineers use the India profile, a generation behind.

A batch summarisation job. Every night the bank summarises the day’s loan files and call transcripts: personal and financial data, high volume, no latency requirement. Here the hosted residency options are thinnest. Azure’s batch discount comes only with deployments that can process outside India. Bedrock lists Mumbai and Hyderabad among the Regions where Claude Opus 5 batch jobs run through a cross-Region profile, which keeps data in India only if the job names the India profile, not the global one.

The better fit is an open-weight model in-Region in Mumbai, where gpt-oss-120b, Qwen3 235B, DeepSeek V3.2 and Mistral Large 3 all support single-Region batch, or the same class of model on the bank’s own GPUs or an Indian GPU cloud. Summarisation is bounded and a queue keeps the accelerators busy: the workload shape that makes self-hosting pay. If the transcripts are in Hindi, Tamil or Marathi, evaluate on them before choosing, because Indic benchmark scores and real traffic diverge.

Three workloads, three answers. The bank needs a router that knows the data class of every request, with three routes behind it: India-only, global and in-house. I set out that router in a reference design in June; residency is one more attribute it routes on.

Recommendations

  1. Write residency requirements per hop, not per vendor. For each data class, state where inference may run, where logs and traces live, where the provider’s safety copy may sit and for how long, and which add-ons are allowed.
  2. Treat APAC as offshore. Bedrock’s apac. profiles and Azure’s Asia Pacific data zone are not Indian residency, and they are priced no lower than global options. Block them for India-only data classes alongside global endpoints: a service control policy on AWS, Azure Policy on deployment types, an organisation policy on Google Cloud.
  3. Prefer options whose safety copy is documented in India, and remove it where you can. Ask for zero retention or modified abuse monitoring where you are eligible, and record the model each workload uses, because retention differs by model behind the same endpoint.
  4. Keep the gateway, logs and tracing in India, and host the tracing yourself. CERT-In already requires the logs to stay here for 180 days. An offshore observability backend undoes residency, whatever the model does.
  5. Plan for the in-country catalogue to trail. Every provider ships new models globally first. Make the India route’s model a configuration value, and re-run the evaluation when a provider adds one, as AWS did for Kimi K3 this week.
  6. Use open weights in-Region or on your own GPUs for batch and bounded work. It is the only option in which every hop is in India by construction, and for batch it is often the cheapest.
  7. Re-verify monthly, from the source. Call GetInferenceProfile from each source Region and diff Microsoft’s region tables and Google’s commitment table against last month’s. The India profile row of Table 1 changed twice between 29 September and today.

In August I argued that the practical way to put frontier-class capability on data that cannot leave India was to host open weights. Today a regulated enterprise can call Claude, GPT-5.6 and Kimi K3 in Mumbai and Hyderabad, an older GPT on reserved capacity in Chennai and a Gemini Flash in Mumbai, and prove where each request ran. That is real progress, and it moves the hard part: the model no longer decides whether your data stayed in India; your gateway, your logs, your tracing and the safety copy you did not read about decide it.

Sources

  1. AWS. Amazon Bedrock User Guide: Cross-Region inference; Supported Regions and models for inference profiles; Regional availability by models; Supported Regions and models for batch inference; model cards for Kimi K3, Claude Opus 5 and Nova Pro; Abuse detection; Document history (India profile for Kimi K3, 6 October 2026). Retrieved October 2026.
  2. AWS Machine Learning Blog. Amazon Bedrock expands Claude model availability to in-country inferencing in India, 29 September 2026; Introducing OpenAI models on Amazon Bedrock for in-country inferencing in India, August 2026.
  3. Microsoft Learn. Understanding deployment types in Microsoft Foundry Models (updated August 2026); Region availability for Foundry Models sold by Azure; Data, privacy, and security for Foundry Models sold by Azure; Abuse monitoring. Microsoft Source Asia, Microsoft’s newest India datacenter region goes live, 6 August 2026.
  4. Google Cloud. Gemini Enterprise Agent Platform documentation: Data residency, Deployments and endpoints, Abuse monitoring and Zero data retention. Last updated 7 October 2026.
  5. Anthropic. Data residency, Claude Platform documentation, retrieved October 2026. Inc42, OpenAI enables local data storage in India, 8 May 2025.
  6. Reserve Bank of India. FAQs: Storage of Payment System Data. 26 June 2019, on the circular of 6 April 2018.
  7. Securities and Exchange Board of India. Framework for Adoption of Cloud Services by SEBI Regulated Entities. 6 March 2023.
  8. CERT-In. Directions under sub-section (6) of section 70B of the Information Technology Act, 2000. 28 April 2022. Cyril Amarchand Mangaldas, Primer on IRDAI Information and Cyber Security Guidelines, 2023, 8 January 2024.
  9. Digital Personal Data Protection Rules, 2025: text of Rule 13 and Rule 15. AZB & Partners, India’s Digital Personal Data Protection Act: phased rollout and key compliance milestones, 14 November 2025.
  10. Press Information Bureau. IndiaAI Mission expands AI ecosystem with affordable compute and startup support. 25 March 2026. Communications Today, IndiaAI Mission faces gap between capacity and deployed GPUs, 1 October 2026. Inc42, Govt assigns AI workloads to Yotta, NxtGen and E2E, 26 April 2025.
  11. Digital Terminal. Yotta launches ₹10,000 free credits program for Shakti Studio AI platform. 8 August 2026. E2E Networks, TIR GenAI API documentation, retrieved October 2026.
Ashish Kumar

Ashish KumarHead of Platforms, AI & Data at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn