On 6 October, AWS added Moonshot’s Kimi K3 to the short list of models that Amazon Bedrock will run only inside India. Requests sent to the new in.moonshotai.kimi-k3 inference profile are processed in Mumbai or Hyderabad and nowhere else. Kimi K3 is the third model family with an India profile, after OpenAI’s GPT-5.6 Terra and Luna in August and Anthropic’s Claude Opus 5, Sonnet 5 and Haiku 4.5 on 29 September. On 7 October, Google updated the data residency page that states where Gemini on Vertex AI processes data. Its India column lists Gemini 3.5 Flash and Gemini 2.5 Flash, and none of the three Flash releases that followed them.
So the question has changed. LLM data residency in India is no longer a matter of whether a hosted model can run in the country: on three hyperscalers and several Indian GPU clouds, it can. The questions now are which models, of which generation, and which hops of a request actually stay. This post is the map, built from the providers’ own documentation as it stood this week. The Claude detail, and what its generation lag costs, is in yesterday’s post. The argument here is broader: residency claims rarely break at the model. They break at the hops around it.
What LLM data residency in India actually covers
Data at rest. Uploaded files, fine-tuned weights, stored conversation state, batch inputs waiting to run. All three hyperscalers keep this in the region or geography you chose, whatever deployment type you pick. It is the easy part, and what most “stored in India” announcements describe.
Inference processing. Where the accelerator that runs the model sits, and so where your prompt and its answer exist in plaintext. This is what inference profiles, deployment types and endpoints change, and where the products differ. A global option on any of the three clouds can process a request wherever the provider chooses. Google is the bluntest: its endpoints “don’t guarantee data residency or in-region ML processing”, and a separate per-model table lists the commitments it does make.
Logs and safety copies. Your own logs, audit trails and traces, which you configure. And the provider’s abuse-monitoring copies, which you do not, and which vary by model. Amazon Bedrock stores no inputs or outputs by default, but keeps classifier-flagged traffic for OpenAI’s GPT-5.6 and GPT-6 models for up to 30 days, and all traffic for Claude Fable 5 and 5.1 for up to 30 days, in the Region that served the request. Azure stores flagged samples for possible human review in your resource’s geography. Google may log flagged prompts for up to 90 days in your project’s region, unless you are on a Google Cloud Master Agreement, which is exempt by default.
Providers sell the second item. Regulators mostly ask about the first and the third. The gap between them is where residency claims fail.
What Indian rules actually require
The common version of the obligation, that Indian law requires AI to run in India, is wrong. No Indian rule says a language model must run in the country. Several say that particular data must be stored, processed or logged here, and they differ by sector.
The DPDP Act and Rules. The Digital Personal Data Protection Act of 2023 permits transfers abroad unless the government restricts a country by notification, and preserves any sectoral law that sets a higher bar. The Rules notified on 13 November 2025 keep that shape. Rule 15 allows transfers subject to requirements the government may set, by order, about making data available to a foreign state. Rule 13(4) is the one to plan for: it lets the government specify, on a committee’s recommendation, personal data that a Significant Data Fiduciary must keep in India, together with the traffic data about its flow. Both come into force in May 2027.
RBI on payment data. The Reserve Bank’s circular of 6 April 2018 requires the entire data relating to payment systems to be stored only in India. Its 2019 FAQ allows processing abroad if the data is deleted there and brought back within one business day or 24 hours, whichever is earlier, and extends the rule to the service providers banks use. A prompt that carries a customer’s transactions is payment data.
SEBI on cloud. SEBI’s cloud framework of 6 March 2023 is the strictest of the set. For exchanges, depositories, brokers, asset managers and its other regulated entities, data in the cloud, logs included, should reside and be processed “within the legal boundaries of India”, in data centres of MeitY-empanelled providers with a valid STQC audit. A hosted model is a cloud service, so sending such data to a model abroad is hard to square with that text.
Insurance. IRDAI’s 2023 guidelines require intermediaries, though not insurers, to keep ICT logs and critical data in India; insurers must keep primary data here under record-keeping rules.
Logs, for everyone. CERT-In’s directions of 28 April 2022 require service providers, data centres, body corporates and government organisations to keep logs of all ICT systems for a rolling 180 days, within the Indian jurisdiction. An LLM gateway is an ICT system.
So the binding constraints attach to data classes, not models: payment data, securities-market data, insurance records and logs, plus a DPDP switch for named categories from 2027. Residency has to be decided per request, not per vendor.
From Mumbai, Bedrock’s APAC profile can route to Tokyo, Seoul, Osaka, Singapore and Sydney. APAC is not India
Amazon Bedrock in Mumbai and Hyderabad: three ways to call a model
AWS has two Indian Regions, Mumbai (ap-south-1) and Hyderabad (ap-south-2), and Bedrock offers three ways to call a model from them. In-Region inference processes the request in the Region you call. A geographic profile routes it among Regions in one geography: the US, the EU, APAC, Japan, Australia or India. A global profile sends it to any commercial Region, about 10 percent cheaper for some models. The model ID says which: a bare ID for in-Region, a prefix such as in. or apac. for a geography, global. for the rest.
This week the India profile carries six models: Claude Opus 5, Sonnet 5 and Haiku 4.5; GPT-5.6 Terra and Luna; and Kimi K3. Each routes only between Mumbai and Hyderabad, and CloudTrail records, written in the Region you called, note which of the two served each request. The newer models are missing. Claude Opus 5.5, Sonnet 5.5 and Fable 5.1, and OpenAI’s GPT-6 Sol, Luna and Astra, reach the Indian Regions only through the global profile.
In-Region inference in Mumbai is a different catalogue, and the strictest residency Bedrock sells. It is almost entirely open-weight: gpt-oss-120b, Qwen3 235B and Qwen3 Coder 480B, DeepSeek V3.2, Mistral Large 3, Kimi K2.5, GLM 5, MiniMax M2.5, Gemma 3 and NVIDIA’s Nemotron Nano family, most of which also run batch jobs in Mumbai alone. Hyderabad’s in-Region list is a single embedding model.
Then there is the trap. Amazon’s Nova Pro, Lite and Micro are offered from Mumbai through a geographic profile, but it is the APAC one, and from Mumbai it routes to Tokyo, Seoul, Osaka, Singapore and Sydney as well. It is priced like any geographic profile and satisfies none of the Indian rules above. AWS’s advice is to call GetInferenceProfile from every source Region and check every destination, rather than rely on “the geographic prefix in an inference profile ID alone”. Enforce it with a service control policy that denies Bedrock inference outside the two Indian Regions: a cross-Region request fails if any destination is blocked, so a pasted global or APAC ID returns an error, not an offshore answer.
Azure OpenAI in India: South India is the region, and the data zone is not India
Microsoft has four cloud regions in India: Central India in Pune, South India in Chennai, West India in Mumbai, and India South Central in Hyderabad, live since 6 August with three availability zones. For Azure OpenAI only one matters today: Microsoft’s availability tables for Foundry Models list South India and no other Indian region.
Azure’s residency choice is the deployment type. Global types may process a prompt in any Azure region. Data Zone types process within the United States, the European Union or Asia Pacific, and Microsoft defines the last as any Asia Pacific nation. Standard and Regional Provisioned deployments process in the deployment’s region, moving within the geography only for operational reasons. Only those two keep inference in India.
What they carry is the question. Standard in South India offers gpt-4.1-mini and gpt-4o, plus embeddings and Whisper. Regional Provisioned, where you reserve throughput units, reaches gpt-5, gpt-5.1 and gpt-5.4-mini. GPT-5.6 and the GPT-6 models are offered from South India only as Global or Asia Pacific Data Zone deployments. Batch has no regional type at all, so a batch deployment in South India is Global Batch and can run anywhere. Microsoft is candid that geography-based types get new models last.
On abuse monitoring, flagged prompts and completions are stored in your resource’s geography and may be reviewed by Microsoft employees. For the European Economic Area the documentation says those reviewers are in the EEA; it says nothing equivalent for India. Eligible customers can apply for modified abuse monitoring, which removes the storage and the human review.
Vertex AI and Gemini in Mumbai: two Flash models and a commitment table
Google Cloud has two Indian regions, Mumbai (asia-south1) and Delhi (asia-south2). For generative AI, Delhi appears in neither the endpoint list nor the residency table, so the map is Mumbai alone. Google, which now documents Vertex AI under its Gemini Enterprise Agent Platform, separates endpoints from commitments: the data residency page states, model by model, where ML processing is guaranteed.
The India column of that table has three entries: Gemini 3.5 Flash, the 128,000-token configuration of Gemini 2.5 Flash, and text-embedding-005. Gemini 3.6, 3.7 and 3.8 Flash carry commitments only for the US and EU multi-regions. No Gemini Pro model has an India commitment. Nor does any partner or open model: Claude on Google Cloud has commitments in Singapore, Belgium or Taiwan for some versions, and none in India.
Two Google features create copies. Grounding with Google Search stores queries derived from prompts for up to three days; if you use it, that storage cannot be switched off, and the documentation does not say where the logs are kept. And Gemini caches inputs and outputs in memory for up to 24 hours by default, within the selected location, which can be turned off per project.

