On 13 November 2025 the Ministry of Electronics and Information Technology published the Digital Personal Data Protection Rules, 2025 in the Gazette as G.S.R. 846(E), with a timetable: the Data Protection Board from that day, Consent Managers after a year, and almost everything a company running AI systems must do after eighteen months. Eight days earlier the ministry had released the India AI Governance Guidelines, and its secretary, S. Krishnan, said the focus remained on “using existing legislation wherever possible”. So the question platform teams keep asking, what India’s AI law requires, has a plain answer: for any system that touches personal data, the DPDP Rules are the AI law. They never mention an LLM or generative AI, and do not need to. The Act defines processing to include storage, retrieval, combination and indexing, which describes a retrieval-augmented LLM application rather well.
Today, 4 October 2026, the eighteen-month date, 13 May 2027, is 221 days away. This is a practical guide to DPDP compliance for AI: what the Rules say and when, how they map onto the places personal data sits in an LLM system, and what to build first. The argument is architectural. Most duties are cheap if personal data lives where you can find, filter, erase and log it, which means the retrieval layer and a log store, and expensive if it lives in model weights, a vendor’s training pipeline or a dashboard nobody owns.
What the DPDP Rules 2025 say, and when each part applies
Rule 1 sets the clock. Rules 1, 2 and 17 to 21, the definitions and the rules that set up the Board as a digital office, came into force on publication. Rule 4, under which Consent Managers register with the Board, commences on 13 November 2026. Rules 3, 5 to 16, 22 and 23 commence on 13 May 2027, and for an AI platform that is nearly everything: notice, security safeguards, breach intimation, retention and erasure, parental consent, the duties of Significant Data Fiduciaries, data principal rights and transfers abroad. A companion notification, G.S.R. 843(E), brings the Act’s sections 3 to 17, which hold every obligation and right, into force on the same date.
The timetable has been questioned once. After meeting platforms in late January 2026, MeitY asked industry for views, by 4 February, on cutting the eighteen months to twelve, which would move the date to 13 November 2026, and on bringing the cross-border rules and the one-year retention duty in Rule 8(3) forward. Officials said no decision had been taken, and as of 4 October no amendment has appeared in the Gazette, so 13 May 2027 stands. The Board exists in law, and on 6 May 2026 MeitY invited applications for its chairperson and four members. Plan for May 2027 with a version that works for November 2026, because the provisions the ministry thought needed no runway, transfers and retention, both reach AI telemetry.
DPDP compliance for AI starts with where personal data sits
I have built conversational systems since 2016: at Active.Ai, where I founded the AI research function and architected the Triniti platform for banks, and at Jio Haptik, which ran conversational AI at scale. Intent discovery, automated evaluation and training-data generation, which my teams worked on at Haptik, all start from the transcript, and the transcript is where customers put whatever personal data they think will solve the problem. LLM systems have multiplied the places it goes.
A support assistant built on an LLM typically sends it to seven: the prompt, holding the customer’s words and anything retrieved or fetched by a tool; the conversation log; the retrieval index; the fine-tuning set; the evaluation set; the model vendor; and analytics. Each is processing in the Act’s sense, with its own owner, retention period and blast radius. The first control, as I argued in classify first, is knowing which of them holds personal data, of what class and for what purpose.
Most DPDP Rules obligations reach AI systems on 13 May 2027, eighteen months after notification
Consent and purpose: what the DPDP Act means for generative AI
Section 4 allows processing on consent or for a legitimate use under section 7. Consent under section 6(1) must be “free, specific, informed, unconditional and unambiguous”, and limited to the data necessary for the specified purpose, which is whatever the notice said. Rule 3 sets out the notice: an itemised description of the data, the purpose and the service it enables, and a link to withdraw consent as easily as it was given, exercise rights and complain to the Board. Section 5(3) lets the person read it in English or any language of the Eighth Schedule, which for an Indic chat product is a design requirement.
For a support assistant the first purpose is easy. A customer who types her order number and complaint into a chat has voluntarily provided personal data for a purpose, and section 7(a) covers processing it for that purpose. The hard purposes come later: fine-tuning on the conversation, building an evaluation set, improving a vendor’s model. None resolves her complaint. If the notice did not name model training, training is a new purpose needing its own notice and consent, and section 6(10) puts the burden of proving both on the fiduciary. That means a consent record keyed to the person, which indexing and training jobs must query.
Two transitional points matter. Section 5(2) covers data collected with consent before the Act commenced: the fiduciary must send a notice as soon as reasonably practicable and may keep processing until consent is withdrawn, which is the position of every ticket archive now being chunked into a vector store. And on withdrawal, section 6(6) requires the fiduciary and its processors to stop within a reasonable time, which in retrieval means removing her chunks from the index, not hiding them in the interface.
Scraped data and the publicly available exemption
Section 3(c)(ii) takes out of the Act personal data made publicly available by the person it relates to, or by someone legally obliged to publish it; the illustration is a person blogging her views on social media. It is often cited to put open-web training outside DPDP, and it is narrower than that. The test is who made the data public, not where you found it. A crawl mixes what people posted about themselves and what the law required to be published, both exempt, with what others posted about them, which a crawler cannot tell apart and the exemption does not cover.
If you buy foundation models, this is first the vendor’s question. It becomes yours when you crawl to build a retrieval corpus, enrich customer records or fine-tune on scraped text: keep provenance per source, and do not assume the exemption survives a join with a record you hold under consent. The research exemption in section 17(2)(b), with standards set by Rule 16 and the Second Schedule, is narrower still, since it excludes data used to take a decision about a specific person. I would not build a product on either.
Children’s data in chat products
A child under the Act is anyone under eighteen. Section 9 requires verifiable parental consent before processing a child’s personal data, forbids processing likely to harm a child’s well-being, and bans tracking, behavioural monitoring and targeted advertising directed at children; a breach can cost up to ₹200 crore. Rule 10 requires the fiduciary to check that the person consenting as parent is an identifiable adult, from identity and age details it holds or that are provided voluntarily, including a virtual token from an authorised entity or details verified through a Digital Locker service provider.
A public chat assistant will have children among its users, intended or not. The Fourth Schedule exempts processing needed to confirm that a user is not a child, and to run the Rule 10 checks, so an age gate at sign-up followed by a parental consent flow is lawful and necessary. Memory is harder. A feature that remembers a child’s interests and personalises suggestions is difficult to distinguish from behavioural monitoring, and the Schedule’s monitoring exemptions belong to named classes such as educational institutions and crèches, not to a chat company selling to families. For child accounts the conservative design is memory off, no personalisation from history, and no advertising.
Erasure belongs in the retrieval layer, not in the weights
Sections 11 to 14 give the data principal the rights a platform must implement: a summary of her data and its processing, with the identities of every fiduciary and processor it was shared with; correction, completion, updating and erasure; grievance redressal; and nomination. Rule 14 caps the period for answering grievances at ninety days. Section 8(7) requires erasure, by the fiduciary and its processors, when consent is withdrawn or the purpose is no longer served. Rule 8 adds a fixed clock for the largest platforms: e-commerce entities and social media intermediaries with at least two crore registered users, and online gaming intermediaries with at least fifty lakh, must erase personal data three years after a user last engaged, with at least 48 hours’ warning.
Now ask where a person lives in an LLM system. In a retrieval index she is a set of chunks carrying her identifier, and erasure is a delete and a re-index. In a fine-tuned model she is spread across the weights, and the only erasure I would defend is retraining without her data; I know of no unlearning method whose guarantee you could put before the Board. The Act does not say whether weights contain personal data, and a platform team need not win that argument if it never has to make it. Keep personal data in retrieval, filtered by the permissions of the person asking and deletable by identifier, and fine-tune only on de-identified or generated data. It is the architecture I argued for on quality grounds in grounding is the product: answers that cite a retrievable source are easier to check, and easier to erase.
The complication is in the Rules themselves. Rule 6(1)(e) requires logs needed to investigate unauthorised access to be kept for a year, and Rule 8(3) requires personal data, traffic data and processing logs to be kept for at least a year for the purposes in the Seventh Schedule, which include lawful requests from the State, and then erased. Erasure yields to retention the law requires. So run two stores on two clocks: a product store, holding the index, memory and anything a model reads, from which people are erased on request; and a restricted log store that no model or analyst reads by default, kept for twelve months and then erased.
Security, logs and the 72-hour breach clock
Failing to keep reasonable security safeguards under section 8(5) is the costliest breach in the Act’s Schedule, up to ₹250 crore. Rule 6 sets the floor: encryption, obfuscation, masking or virtual tokens mapped to personal data; access control, including over the processor’s systems; logs and monitoring able to detect unauthorised access; backups; a year of logs; and a security clause in the processor contract. For an LLM, the first item is an instruction: tokenise identifiers before they enter a prompt, and resolve them only where the real value is needed.
Section 2(u) treats any unauthorised processing, or accidental disclosure or loss of access, that compromises confidentiality, integrity or availability as a personal data breach. A retriever that shows one customer another customer’s ticket has disclosed personal data; so has an agent that obeys instructions hidden in a document and emails a record outside the company; so has a vendor whose prompt logs are exposed. Each starts the Rule 7 clock: every affected person told without delay what happened, the likely consequences and what she can do; the Board told without delay and, within 72 hours of the fiduciary becoming aware, given a detailed account of facts, causes, mitigation and remediation. Failing to notify carries up to ₹200 crore.
I set out one incident clock for the DPDP Rules, the AI Act and the RBI in two regulators, seven controls. For an LLM system the addition is detection: a cross-customer leak looks like a normal answer unless something checks the owner of every retrieved chunk against the identity of the person asking.
Significant Data Fiduciaries and the algorithmic software duty
Section 10 lets the government designate fiduciaries as significant on factors from the volume and sensitivity of data to electoral democracy and public order. It commences in May 2027 and no one has yet been designated, but large banks, insurers, platforms and telecom operators should plan as if they will be. A Significant Data Fiduciary needs a Data Protection Officer based in India and answerable to its board, an independent data auditor, and periodic impact assessments and audits; breaches can cost up to ₹150 crore.
Rule 13 makes that a yearly cycle, with a report of significant observations to the Board, and adds two duties. Rule 13(3) requires due diligence that technical measures, including algorithmic software, used for hosting, display, transmission, storage or sharing of personal data are “not likely to pose a risk to the rights of Data Principals”. Rule 13(4) lets the government, on a committee’s recommendation, require specified personal data and the traffic data about its flow to stay in India.
Rule 13(3) is the closest Indian law comes to an AI audit duty, and its verbs fit an LLM application, which displays, transmits, stores and shares personal data on every turn. Whether a model is likely to pose a risk is an evaluation question. Test that retrieval never returns another customer’s data, that the model does not reproduce personal data from its fine-tuning set, that injected instructions cannot trigger a disclosure, and that outputs are accurate where section 8(3) requires it because a decision will follow. Run the tests on every model change, as part of the bench described in the evaluation bench, in full, and the annual assessment summarises results you already hold.
Model vendors are processors, and transfers abroad are transfers
A model provider that processes prompts on your behalf is a Data Processor, and section 8(1) leaves the fiduciary responsible for what it does. Section 8(2) requires a valid contract, Rule 6(1)(f) a security clause in it, and sections 6(6) and 8(7)(b) require the fiduciary to make the processor stop and erase when consent is withdrawn or the purpose ends. An access request under section 11 must name the processors, so your model vendors appear in the answer.
The clause that matters most is the vendor’s own use. A company that trains its own models on your customers’ prompts has chosen a purpose of its own, and choosing the purpose is what makes someone a fiduciary under section 2(i). Contract for no training, short retention, deletion on request, disclosure of sub-processors, and breach notice fast enough to leave room inside your 72 hours.
Transfers abroad are permitted unless the government restricts them by notification under section 16(1), or sets requirements under Rule 15 on data reaching a foreign State. Two things narrow that: section 16(2) preserves stricter laws, so sectoral rules such as the RBI’s on payment data still govern prompts containing such data, and Rule 13(4) can keep specified data of a Significant Data Fiduciary in India. A prompt sent to an endpoint abroad is a transfer, and so is a trace shipped to an observability service in another region.
A worked example: one support agent, seven data flows
Take an e-commerce company with three crore registered users that deploys a support agent: an LLM from a hosted vendor, retrieval over five years of past tickets, tool calls to the order system, and a plan to fine-tune a smaller model on its best conversations. With more than two crore users it falls in the Third Schedule, and it is a plausible Significant Data Fiduciary.
Prompt. A customer’s message, her order history from a tool call and three retrieved chunks form the prompt. Support processing rests on section 7(a), and the Rule 3 notice should say that an AI assistant reads her messages. Under Rule 6(1)(a), phone numbers, addresses and payment references become tokens before the prompt leaves the company, and the tool layer, not the model, resolves them.
Conversation log. Written to the restricted log store, access-logged, closed to training and analytics, and erased after twelve months under Rules 6(1)(e) and 8(3).
Retrieval index. Mostly tickets collected before the Act, so section 5(2) requires a notice to those customers. Each chunk carries the customer’s identifier and a permission tag, and the retriever filters by the identity of the person asking; an agent acting for her carries her scope, as in an agent is a principal. Erasure deletes by identifier and re-indexes, and the three-year inactivity clock runs against the same identifiers.
Fine-tuning set. A new purpose: either a specific consent enforced as a filter or, better, de-identified and synthetic conversations only, so that no erasure request can reach the weights.
Evaluation set. A few thousand de-identified conversations with a named owner, refreshed rather than kept indefinitely. The leakage and injection tests that evidence Rule 13(3) due diligence live here.
Model vendor. A processor under contract, named in access responses. An endpoint abroad is lawful unless the government restricts it, and the router keeps sectorally restricted data on an endpoint in India.
Analytics. Sentiment scores and dashboards are a separate purpose built on the log: aggregate them, strip identifiers and exclude child accounts from profiling.
Then the drill: retrieval returns another customer’s order. Detection fires because the chunk’s owner does not match the requester; the customer whose order leaked is told without delay; the Board gets a first description at once and the full report within 72 hours; the root cause, a missing filter on one index, goes into the Rule 13 audit. Table 1 sets out the seven flows and the breach path.