Built for the Real World · Essay · Governance

DPDP Rules for AI systems: what the 18-month clock asks of your models, prompts and logs

The DPDP Rules 2025 were notified on 13 November 2025, and most of their duties bind on 13 May 2027. India chose no separate AI law, so for any AI system that touches personal data these Rules are the AI law. Here is what they ask of an LLM application’s prompts, retrieval indexes, fine-tuning sets, logs and vendors, and what to build first.

Abstract illustration for “DPDP Rules for AI systems: what the 18-month clock asks of your models, prompts and logs”

On 13 November 2025 the Ministry of Electronics and Information Technology published the Digital Personal Data Protection Rules, 2025 in the Gazette as G.S.R. 846(E), with a timetable: the Data Protection Board from that day, Consent Managers after a year, and almost everything a company running AI systems must do after eighteen months. Eight days earlier the ministry had released the India AI Governance Guidelines, and its secretary, S. Krishnan, said the focus remained on “using existing legislation wherever possible”. So the question platform teams keep asking, what India’s AI law requires, has a plain answer: for any system that touches personal data, the DPDP Rules are the AI law. They never mention an LLM or generative AI, and do not need to. The Act defines processing to include storage, retrieval, combination and indexing, which describes a retrieval-augmented LLM application rather well.

Today, 4 October 2026, the eighteen-month date, 13 May 2027, is 221 days away. This is a practical guide to DPDP compliance for AI: what the Rules say and when, how they map onto the places personal data sits in an LLM system, and what to build first. The argument is architectural. Most duties are cheap if personal data lives where you can find, filter, erase and log it, which means the retrieval layer and a log store, and expensive if it lives in model weights, a vendor’s training pipeline or a dashboard nobody owns.

What the DPDP Rules 2025 say, and when each part applies

Rule 1 sets the clock. Rules 1, 2 and 17 to 21, the definitions and the rules that set up the Board as a digital office, came into force on publication. Rule 4, under which Consent Managers register with the Board, commences on 13 November 2026. Rules 3, 5 to 16, 22 and 23 commence on 13 May 2027, and for an AI platform that is nearly everything: notice, security safeguards, breach intimation, retention and erasure, parental consent, the duties of Significant Data Fiduciaries, data principal rights and transfers abroad. A companion notification, G.S.R. 843(E), brings the Act’s sections 3 to 17, which hold every obligation and right, into force on the same date.

The timetable has been questioned once. After meeting platforms in late January 2026, MeitY asked industry for views, by 4 February, on cutting the eighteen months to twelve, which would move the date to 13 November 2026, and on bringing the cross-border rules and the one-year retention duty in Rule 8(3) forward. Officials said no decision had been taken, and as of 4 October no amendment has appeared in the Gazette, so 13 May 2027 stands. The Board exists in law, and on 6 May 2026 MeitY invited applications for its chairperson and four members. Plan for May 2027 with a version that works for November 2026, because the provisions the ministry thought needed no runway, transfers and retention, both reach AI telemetry.

DPDP Rules 2025 commencement timelineA horizontal timeline from November 2025 to May 2027. Three commencement dates sit on the axis: 13 November 2025, when Rules 1, 2 and 17 to 21 came into force and the Data Protection Board was set up; 13 November 2026, when Rule 4 on Consent Managers commences; and 13 May 2027, highlighted, when Rules 3, 5 to 16, 22 and 23 commence, covering notice, security, breach, erasure, children, Significant Data Fiduciaries, rights and transfers. Below the axis are three context markers: the January 2026 consultation on shortening the timeline, which was not notified; the 6 May 2026 call for Board members; and 4 October 2026, the date of this post, 40 days before Rule 4 and 221 days before 13 May 2027. THREE COMMENCEMENT DATES UNDER RULE 1G.S.R. 846(E), GAZETTE OF 13 NOV 2025 13 Nov 2025 · in force Rules 1, 2 and 17 to 21 Definitions, and the rules that set up the Data Protection Board as a digital office 13 Nov 2026 · after one year Rule 4 Consent Managers register with the Board; Act s. 6(9) and s. 27(1)(d) commence 13 May 2027 · after 18 months Rules 3, 5 to 16, 22 and 23 Notice, security, breach, erasure, children, SDF duties, rights and transfers abroad 18-month runway Nov 2025Feb 2026May 2026Aug 2026Nov 2026Feb 2027May 2027 Late Jan 2026 MeitY consults on cutting 18 months to 12; not notified 6 May 2026 MeitY invites applications for a Board chairperson and four members 4 Oct 2026 · this post 40 days to Rule 4 221 days to 13 May 2027
Figure 1. The DPDP Rules commence in three steps under Rule 1. The obligations that reach an AI system, from notice to breach reporting to Significant Data Fiduciary duties, all arrive on 13 May 2027; the January 2026 proposal to shorten the runway was not notified. Sources: G.S.R. 846(E); PIB, 17 November 2025; MeitY, 6 May 2026; Storyboard18, 28 January 2026.

DPDP compliance for AI starts with where personal data sits

I have built conversational systems since 2016: at Active.Ai, where I founded the AI research function and architected the Triniti platform for banks, and at Jio Haptik, which ran conversational AI at scale. Intent discovery, automated evaluation and training-data generation, which my teams worked on at Haptik, all start from the transcript, and the transcript is where customers put whatever personal data they think will solve the problem. LLM systems have multiplied the places it goes.

A support assistant built on an LLM typically sends it to seven: the prompt, holding the customer’s words and anything retrieved or fetched by a tool; the conversation log; the retrieval index; the fine-tuning set; the evaluation set; the model vendor; and analytics. Each is processing in the Act’s sense, with its own owner, retention period and blast radius. The first control, as I argued in classify first, is knowing which of them holds personal data, of what class and for what purpose.

Most DPDP Rules obligations reach AI systems on 13 May 2027, eighteen months after notification

Consent and purpose: what the DPDP Act means for generative AI

Section 4 allows processing on consent or for a legitimate use under section 7. Consent under section 6(1) must be “free, specific, informed, unconditional and unambiguous”, and limited to the data necessary for the specified purpose, which is whatever the notice said. Rule 3 sets out the notice: an itemised description of the data, the purpose and the service it enables, and a link to withdraw consent as easily as it was given, exercise rights and complain to the Board. Section 5(3) lets the person read it in English or any language of the Eighth Schedule, which for an Indic chat product is a design requirement.

For a support assistant the first purpose is easy. A customer who types her order number and complaint into a chat has voluntarily provided personal data for a purpose, and section 7(a) covers processing it for that purpose. The hard purposes come later: fine-tuning on the conversation, building an evaluation set, improving a vendor’s model. None resolves her complaint. If the notice did not name model training, training is a new purpose needing its own notice and consent, and section 6(10) puts the burden of proving both on the fiduciary. That means a consent record keyed to the person, which indexing and training jobs must query.

Two transitional points matter. Section 5(2) covers data collected with consent before the Act commenced: the fiduciary must send a notice as soon as reasonably practicable and may keep processing until consent is withdrawn, which is the position of every ticket archive now being chunked into a vector store. And on withdrawal, section 6(6) requires the fiduciary and its processors to stop within a reasonable time, which in retrieval means removing her chunks from the index, not hiding them in the interface.

Scraped data and the publicly available exemption

Section 3(c)(ii) takes out of the Act personal data made publicly available by the person it relates to, or by someone legally obliged to publish it; the illustration is a person blogging her views on social media. It is often cited to put open-web training outside DPDP, and it is narrower than that. The test is who made the data public, not where you found it. A crawl mixes what people posted about themselves and what the law required to be published, both exempt, with what others posted about them, which a crawler cannot tell apart and the exemption does not cover.

If you buy foundation models, this is first the vendor’s question. It becomes yours when you crawl to build a retrieval corpus, enrich customer records or fine-tune on scraped text: keep provenance per source, and do not assume the exemption survives a join with a record you hold under consent. The research exemption in section 17(2)(b), with standards set by Rule 16 and the Second Schedule, is narrower still, since it excludes data used to take a decision about a specific person. I would not build a product on either.

Children’s data in chat products

A child under the Act is anyone under eighteen. Section 9 requires verifiable parental consent before processing a child’s personal data, forbids processing likely to harm a child’s well-being, and bans tracking, behavioural monitoring and targeted advertising directed at children; a breach can cost up to ₹200 crore. Rule 10 requires the fiduciary to check that the person consenting as parent is an identifiable adult, from identity and age details it holds or that are provided voluntarily, including a virtual token from an authorised entity or details verified through a Digital Locker service provider.

A public chat assistant will have children among its users, intended or not. The Fourth Schedule exempts processing needed to confirm that a user is not a child, and to run the Rule 10 checks, so an age gate at sign-up followed by a parental consent flow is lawful and necessary. Memory is harder. A feature that remembers a child’s interests and personalises suggestions is difficult to distinguish from behavioural monitoring, and the Schedule’s monitoring exemptions belong to named classes such as educational institutions and crèches, not to a chat company selling to families. For child accounts the conservative design is memory off, no personalisation from history, and no advertising.

Erasure belongs in the retrieval layer, not in the weights

Sections 11 to 14 give the data principal the rights a platform must implement: a summary of her data and its processing, with the identities of every fiduciary and processor it was shared with; correction, completion, updating and erasure; grievance redressal; and nomination. Rule 14 caps the period for answering grievances at ninety days. Section 8(7) requires erasure, by the fiduciary and its processors, when consent is withdrawn or the purpose is no longer served. Rule 8 adds a fixed clock for the largest platforms: e-commerce entities and social media intermediaries with at least two crore registered users, and online gaming intermediaries with at least fifty lakh, must erase personal data three years after a user last engaged, with at least 48 hours’ warning.

Now ask where a person lives in an LLM system. In a retrieval index she is a set of chunks carrying her identifier, and erasure is a delete and a re-index. In a fine-tuned model she is spread across the weights, and the only erasure I would defend is retraining without her data; I know of no unlearning method whose guarantee you could put before the Board. The Act does not say whether weights contain personal data, and a platform team need not win that argument if it never has to make it. Keep personal data in retrieval, filtered by the permissions of the person asking and deletable by identifier, and fine-tune only on de-identified or generated data. It is the architecture I argued for on quality grounds in grounding is the product: answers that cite a retrievable source are easier to check, and easier to erase.

The complication is in the Rules themselves. Rule 6(1)(e) requires logs needed to investigate unauthorised access to be kept for a year, and Rule 8(3) requires personal data, traffic data and processing logs to be kept for at least a year for the purposes in the Seventh Schedule, which include lawful requests from the State, and then erased. Erasure yields to retention the law requires. So run two stores on two clocks: a product store, holding the index, memory and anything a model reads, from which people are erased on request; and a restricted log store that no model or analyst reads by default, kept for twelve months and then erased.

Security, logs and the 72-hour breach clock

Failing to keep reasonable security safeguards under section 8(5) is the costliest breach in the Act’s Schedule, up to ₹250 crore. Rule 6 sets the floor: encryption, obfuscation, masking or virtual tokens mapped to personal data; access control, including over the processor’s systems; logs and monitoring able to detect unauthorised access; backups; a year of logs; and a security clause in the processor contract. For an LLM, the first item is an instruction: tokenise identifiers before they enter a prompt, and resolve them only where the real value is needed.

Section 2(u) treats any unauthorised processing, or accidental disclosure or loss of access, that compromises confidentiality, integrity or availability as a personal data breach. A retriever that shows one customer another customer’s ticket has disclosed personal data; so has an agent that obeys instructions hidden in a document and emails a record outside the company; so has a vendor whose prompt logs are exposed. Each starts the Rule 7 clock: every affected person told without delay what happened, the likely consequences and what she can do; the Board told without delay and, within 72 hours of the fiduciary becoming aware, given a detailed account of facts, causes, mitigation and remediation. Failing to notify carries up to ₹200 crore.

I set out one incident clock for the DPDP Rules, the AI Act and the RBI in two regulators, seven controls. For an LLM system the addition is detection: a cross-customer leak looks like a normal answer unless something checks the owner of every retrieved chunk against the identity of the person asking.

Significant Data Fiduciaries and the algorithmic software duty

Section 10 lets the government designate fiduciaries as significant on factors from the volume and sensitivity of data to electoral democracy and public order. It commences in May 2027 and no one has yet been designated, but large banks, insurers, platforms and telecom operators should plan as if they will be. A Significant Data Fiduciary needs a Data Protection Officer based in India and answerable to its board, an independent data auditor, and periodic impact assessments and audits; breaches can cost up to ₹150 crore.

Rule 13 makes that a yearly cycle, with a report of significant observations to the Board, and adds two duties. Rule 13(3) requires due diligence that technical measures, including algorithmic software, used for hosting, display, transmission, storage or sharing of personal data are “not likely to pose a risk to the rights of Data Principals”. Rule 13(4) lets the government, on a committee’s recommendation, require specified personal data and the traffic data about its flow to stay in India.

Rule 13(3) is the closest Indian law comes to an AI audit duty, and its verbs fit an LLM application, which displays, transmits, stores and shares personal data on every turn. Whether a model is likely to pose a risk is an evaluation question. Test that retrieval never returns another customer’s data, that the model does not reproduce personal data from its fine-tuning set, that injected instructions cannot trigger a disclosure, and that outputs are accurate where section 8(3) requires it because a decision will follow. Run the tests on every model change, as part of the bench described in the evaluation bench, in full, and the annual assessment summarises results you already hold.

Model vendors are processors, and transfers abroad are transfers

A model provider that processes prompts on your behalf is a Data Processor, and section 8(1) leaves the fiduciary responsible for what it does. Section 8(2) requires a valid contract, Rule 6(1)(f) a security clause in it, and sections 6(6) and 8(7)(b) require the fiduciary to make the processor stop and erase when consent is withdrawn or the purpose ends. An access request under section 11 must name the processors, so your model vendors appear in the answer.

The clause that matters most is the vendor’s own use. A company that trains its own models on your customers’ prompts has chosen a purpose of its own, and choosing the purpose is what makes someone a fiduciary under section 2(i). Contract for no training, short retention, deletion on request, disclosure of sub-processors, and breach notice fast enough to leave room inside your 72 hours.

Transfers abroad are permitted unless the government restricts them by notification under section 16(1), or sets requirements under Rule 15 on data reaching a foreign State. Two things narrow that: section 16(2) preserves stricter laws, so sectoral rules such as the RBI’s on payment data still govern prompts containing such data, and Rule 13(4) can keep specified data of a Significant Data Fiduciary in India. A prompt sent to an endpoint abroad is a transfer, and so is a trace shipped to an observability service in another region.

A worked example: one support agent, seven data flows

Take an e-commerce company with three crore registered users that deploys a support agent: an LLM from a hosted vendor, retrieval over five years of past tickets, tool calls to the order system, and a plan to fine-tune a smaller model on its best conversations. With more than two crore users it falls in the Third Schedule, and it is a plausible Significant Data Fiduciary.

Prompt. A customer’s message, her order history from a tool call and three retrieved chunks form the prompt. Support processing rests on section 7(a), and the Rule 3 notice should say that an AI assistant reads her messages. Under Rule 6(1)(a), phone numbers, addresses and payment references become tokens before the prompt leaves the company, and the tool layer, not the model, resolves them.

Conversation log. Written to the restricted log store, access-logged, closed to training and analytics, and erased after twelve months under Rules 6(1)(e) and 8(3).

Retrieval index. Mostly tickets collected before the Act, so section 5(2) requires a notice to those customers. Each chunk carries the customer’s identifier and a permission tag, and the retriever filters by the identity of the person asking; an agent acting for her carries her scope, as in an agent is a principal. Erasure deletes by identifier and re-indexes, and the three-year inactivity clock runs against the same identifiers.

Fine-tuning set. A new purpose: either a specific consent enforced as a filter or, better, de-identified and synthetic conversations only, so that no erasure request can reach the weights.

Evaluation set. A few thousand de-identified conversations with a named owner, refreshed rather than kept indefinitely. The leakage and injection tests that evidence Rule 13(3) due diligence live here.

Model vendor. A processor under contract, named in access responses. An endpoint abroad is lawful unless the government restricts it, and the router keeps sectorally restricted data on an endpoint in India.

Analytics. Sentiment scores and dashboards are a separate purpose built on the log: aggregate them, strip identifiers and exclude child accounts from profiling.

Then the drill: retrieval returns another customer’s order. Detection fires because the chunk’s owner does not match the requester; the customer whose order leaked is told without delay; the Board gets a first description at once and the full report within 72 hours; the root cause, a missing filter on one index, goes into the Rule 13 audit. Table 1 sets out the seven flows and the breach path.

Data flowPersonal data it holdsWhat the DPDP framework requiresAct section or RuleControl to build
Prompt
assembled per model call
Customer’s words, order history from tool calls, retrieved chunksNotice that names the AI use; processing limited to the support purpose; security safeguardss. 5, 6, 7(a); Rules 3, 6(1)(a)Tokenise identifiers before the prompt; resolve them in the tool layer, not the model
Conversation log
transcript and tool calls
Full conversation, tool results, model outputsLogs able to detect and investigate unauthorised access; at least one year of retention, then erasureRules 6(1)(c), 6(1)(e), 8(3)Restricted log store, access-logged, twelve-month clock, closed to training and analytics
Retrieval index
vector store of past tickets
Chunks keyed to individual customersNotice for data collected earlier; stop and erase on withdrawal or request; inactivity erasure for large platforms; no cross-customer disclosures. 5(2), 6(6), 8(7), 12(3); Rules 8(1), 8(2)Identity-aware retrieval filter; delete by identifier and re-index; provenance check on every chunk
Fine-tuning set
examples for adaptation
Selected conversationsA new purpose needs its own notice and consent, provable by the fiduciary; erasure must stay possibles. 6(1), 6(10), 8(7)De-identified or synthetic data only; a consent filter if personal data is unavoidable
Evaluation set
release and regression tests
Sampled production conversationsPurpose limitation; accuracy where decisions follow; evidence of algorithmic due diligence for SDFss. 8(3); Rule 13(3)Small, de-identified, owned and refreshed; leakage and injection tests on every release
Model vendor
Data Processor
Prompts and outputs in transit; sometimes vendor-side logsFiduciary liable for the processor; valid contract with security terms; cessation and erasure; transfer restrictionss. 8(1), 8(2), 16; Rules 6(1)(f), 15, 13(4)No-training clause, short retention, deletion, breach notice in hours; restricted data routed to an endpoint in India
Analytics and telemetry
dashboards, scores, traces
Sentiment scores, resolution data, tracesPurpose limitation; no tracking or behavioural monitoring of children; traces abroad are transferss. 9(3), 16; Rule 13(4)Aggregate and strip identifiers; exclude child accounts; keep unmasked traces in region
Any flow, on a leak
the breach path
Whatever was disclosedTell each affected person without delay; tell the Board without delay and in detail within 72 hourss. 2(u), 8(6); Rule 7Provenance-based detection wired to an incident clock, drilled before May 2027
Table 1. The seven data flows of an LLM support agent with retrieval over past tickets, plus the breach path, mapped to the DPDP obligation each triggers and the control that meets it. Section numbers are from the Digital Personal Data Protection Act, 2023; rule numbers from G.S.R. 846(E) of 13 November 2025. Almost all of these obligations commence on 13 May 2027.

Recommendations before 13 May 2027

  1. Map the seven flows for every AI system, with each one’s owner, data classes, purpose and retention.
  2. Rewrite the notice for AI, and send the legacy notice. Name the AI use, itemise the data, offer the Eighth Schedule languages and link to withdrawal; then send the section 5(2) notice to customers whose historical data sits in retrieval indexes.
  3. Make consent machine-readable. Keep a record keyed to the person that indexing, training and evaluation jobs must query, so that a purpose the customer did not agree to is blocked by a filter, not a policy.
  4. Keep personal data out of the weights. Fine-tune on de-identified or generated data; keep personal data in retrieval, filtered by the identity of the person asking and deletable by identifier well inside ninety days, and keep logs in a separate store on a twelve-month clock.
  5. Treat a cross-customer answer as a breach. Tokenise identifiers before the prompt, test retrieval isolation and injection on every release, and wire leak detection to a drilled 72-hour clock.
  6. Age-gate chat products, run the Rule 10 parental consent flow, and switch off memory, personalisation and advertising on children’s accounts.
  7. Re-paper every model and observability vendor as a processor on the contract terms above, and list which endpoints sit outside India.
  8. If you could be designated significant, start the Rule 13 cycle now, with an India-based Data Protection Officer and a first impact assessment written from evaluation results you already have.

The DPDP Rules were not written for language models, and that is their advantage for a platform team. They ask for what a well-built AI system produces as a by-product: a record of where each person’s data is, a way to remove it, a log kept for a year and a clock that starts when something leaks. Two hundred and twenty-one days is enough to build that, if the work starts with the data path rather than the policy.

Sources

  1. Ministry of Electronics and Information Technology. Digital Personal Data Protection Rules, 2025, notification G.S.R. 846(E). Gazette of India, 13 November 2025.
  2. Ministry of Electronics and Information Technology. The Digital Personal Data Protection Act, 2023 (22 of 2023). Gazette of India, 11 August 2023.
  3. Press Information Bureau. DPDP Rules, 2025 notified: a citizen-centric framework for privacy protection. 17 November 2025.
  4. Taxmann. Govt notifies commencement dates for DPDP Act, 2023 (G.S.R. 843(E)). 15 November 2025.
  5. Press Information Bureau. MeitY unveils India AI Governance Guidelines under IndiaAI Mission. 5 November 2025.
  6. Storyboard18. MeitY seeks industry views on fast-tracking DPDP Act rollout, proposes 12-month compliance timeline. 28 January 2026.
  7. S.S. Rana & Co. MeitY plans to cut short DPDP compliance timeline and notify cross-border restrictions for SDFs. 13 February 2026.
  8. Ministry of Electronics and Information Technology. Appointment to the post of Chairperson and other Members in the Data Protection Board of India. 6 May 2026.
Ashish Kumar

Ashish KumarHead of Platforms, AI & Data at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn