Built for the Real World · Essay · Platform design

Classify first: the control every other AI control depends on

The NCSC’s interim guidance of 20 August opens by asking what could go wrong, which nobody can answer without knowing what data an agent can reach. Article 50, the DPDP Rules and the RBI’s localisation rules all presume the same answer. Classification is the control every other AI control reads, and the one almost nobody budgets for. Here is how to build it.

Abstract illustration for “Classify first: the control every other AI control depends on”

On 20 August 2026 the UK’s National Cyber Security Centre published interim guidance on managing the cyber risk of agentic AI, and its first consideration is to identify what could go wrong: document what the agent is meant to do, then threat-model its instructions and “the networks and services the agent may be able to access”. It has a prerequisite the guidance does not state: you cannot say what could go wrong with an agent until you know what data it can reach, and most organisations cannot say that. Eighteen days earlier, on 2 August, Article 50 of the EU AI Act began to apply: a person interacting with an AI system must be told so, and synthetic content must carry a machine-readable mark, with a grace period to 2 December 2026 for systems already on the market. Both duties assume the pipeline knows which outputs are synthetic and which surfaces reach a natural person. In India, the Digital Personal Data Protection Rules, notified on 13 November 2025, allow the Central Government, on the recommendation of a committee, to specify categories of personal data that a Significant Data Fiduciary must keep inside the country, together with the traffic data about its flow. Beneath them sits the Reserve Bank of India’s circular of 6 April 2018, which requires the entire data relating to payment systems to be stored only in India, and the RBI’s FAQ of 26 June 2019, which allows processing abroad provided the data is deleted there within one business day or 24 hours of processing, whichever is earlier. Four documents, three jurisdictions, and each presumes an answer to the same prior question: what class of data is this?

Classification is the control every other AI control depends on, and the one almost nobody budgets for. Routing, residency, de-identification, retention, what an agent may read and what may leave the building are all decisions taken on a class, and when the class is missing each control quietly runs on a default, usually “internal”, the class that lets everything through. This post builds classification as a platform component rather than a policy document, ending with one customer email walked through an agent. Last week I argued that a platform team should build the agent controls three regulators have converged on once, as a control plane. Classification is the layer beneath it; build it second and the plane runs on guesses.

Every control is downstream

Ask what each existing control reads. The router I described in June carries a data-sensitivity field in every task-class entry, and a restricted class never falls back to a cloud tier. That field is a declaration by the developer who wrote the step: a ceiling, correct only if the payload matches it, and a step declared internal that receives a pasted bank statement is routed as internal. Residency is the same decision with geography attached: an endpoint inside India satisfies the RBI only for calls that carry payments data, and a platform that cannot tell which those are either sends everything to the expensive local tier or sends payments data abroad and hopes.

De-identification can only remove what it detects, and detection is classification at the span level. Retention is a per-record decision: twelve months of traces is the right setting for audit evidence and the wrong one for a prompt that contained a one-time password. The NCSC asks that agent logs be kept and monitored as part of security operations, the DPDP Rules for one year of logs after a breach, the RBI that payment data be stored only in India; one trace must satisfy all three, which is possible only if it knows what it contains.

What an agent may read is an entitlement, and I argued in July that entitlements belong to a principal in the directory, not to a prompt. An entitlement is a grant over a class of data, meaningless if the connector behind it returns rows of mixed class, which every real database does. What may leave the building is an egress rule written per destination, and a destination allowed for public data is not allowed for regulated data, so the rule must be indexed by class. Even retrieval returns passages whose class the generation step must inherit, and a retriever that drops the class on the way back has laundered the document.

The dependency runs one way: none of these controls produces a class and all of them consume one. If the class is absent they do not fail. They succeed on a default, which is worse, because the evidence says a control ran.

A scheme that fits models

A scheme for AI is decided by software on every call, on fragments rather than files, assembled from different sources. Four classes are enough; the sub-tags do the regulatory work.

Public. Content the organisation has published or would publish. Any tier, any region, retention as convenient, and Article 50’s mark if the content is synthetic.

Internal. Business information that is neither personal nor published. Contracted cloud tiers with no training on inputs, the regions named in the contract, traces kept for the audit period.

Confidential. Personal data, source code, customer transcripts, pricing, anything under a non-disclosure agreement. Zero-retention contracted tiers or in-boundary tiers; erasure once the purpose is served, as the DPDP Act requires; the 72-hour breach report and the one-year log duty as the clocks.

Regulated. Data a sector regulator has named: the RBI’s payment system data, which its FAQ enumerates as customer data, payment-sensitive data, payment credentials and transaction data; and, once the committee reports, the categories the Central Government specifies for Significant Data Fiduciaries under Rule 12 of the DPDP Rules. Tiers inside the organisation’s boundary only: open weights on hardware it controls, or a contracted endpoint whose region, sub-processors and deletion are in writing. Stored in India; processed abroad only under the RBI’s 24-hour deletion rule, which no platform should rely on for a call whose provider-side logs it cannot inspect.

Four sub-tags cut across the classes: personal, which triggers the DPDP duties; payments, which triggers the RBI’s; code, which triggers the licence and secrets checks; and transcript, which marks a customer’s own words, personal data and the subject of Article 50 when a reply goes back. A pull request embedding a test fixture with real customer rows is code and personal, which is how source code becomes regulated without anyone deciding it should.

Two rules make the scheme usable by a machine. The first is the high-water rule: a prompt inherits the highest class of anything in it, and the union of the sub-tags. A context built from a public passage, an internal runbook and one confidential email is confidential. The second is that class goes down only through a named operation that produces a new object with its own record: masking, aggregation, redaction, or summarisation by a tier allowed for the source class, each a pipeline step with an input hash, an output hash, a method and an owner, never a side effect of formatting. Without the second rule the first is a tax every team evades by copying the text into a new variable.

ClassExamplesAllowed model tiersResidencyRetentionDriving regulation
Public
no sub-tags
Marketing copy, open documentation, filings, press releasesAny tier, any providerAny regionAs convenientAI Act Article 50: machine-readable mark if synthetic and published (from 2 August 2026)
Internal
code without secrets or customer rows
Roadmaps, wikis, meeting notes, non-customer analytics, most source codeContracted cloud tiers, enterprise terms, no training on inputsContracted regionsAudit period for traces (twelve months)Contract and trade secrecy; no sector regulator
Confidential
personal · code · transcript
Customer records, support transcripts, pricing, NDA material, code with secretsZero-retention contracted tiers, or in-boundary tiersIndia for personal data of Indian residentsPurpose-limited; one year of logs after a breachDPDP Act and Rules 2025: 72-hour breach report, annual DPIA and audit for Significant Data Fiduciaries
Regulated
payments · specified categories · health · credit
Payment credentials, transaction data, OTPs, PINs, account details; statements; KYC files; categories the Central Government specifiesIn-boundary tiers only: open weights on own hardware, or an endpoint in India with region, sub-processors and deletion in writingIndia only; abroad only with deletion within 24 h or one business dayThe regulator’s clock; stored only in IndiaRBI circular of 6 April 2018 and FAQ of 26 June 2019; DPDP Rule 12 for Significant Data Fiduciaries
Table 1. The four classes, the sub-tags that carry the regulatory logic, the model tiers each may use, where it may be processed and for how long it may be kept. Sources: RBI circular DPSS.CO.OD No.2785/06.08.005/2017-2018 and FAQ of 26 June 2019; DPDP Rules 2025 as summarised by DLA Piper and SFLC.in; artificialintelligenceact.eu, Article 50.

Classification at the point of use

A scheme is applied in two places. The first is ingestion. When a connector pulls a document, a row, a ticket or a recording into the platform, it attaches the class and sub-tags the source system already knows: the core banking system knows its ledger is payments data. Provenance is the cheapest and most reliable signal there is, and it is lost the moment data is copied out of the system that knew it. The tag travels with the record into the index, the cache and the memory store, and comes back with every passage.

The second place is the gateway, where the prompt is assembled. Provenance says nothing about the free text the user typed, the attachment they dropped in, or the rows a tool returned a second ago. A content classifier covers that: the gateway runs it over every fragment of every call, takes the highest class across provenance and content, and routes accordingly.

A small model can decide more in one pass than most teams assume. Open encoder classifiers, the family I have published on Hugging Face for other purposes, return a fixed set of labels with scores in one forward pass. Fine-tuned on labelled fragments from the organisation’s own traffic, one can decide whether a fragment is a customer’s own words, contains personal identifiers, contains source code, or reads like a financial record, and emit all four. It cannot decide whether a particular name belongs to a customer, and should not try. Deterministic detectors sit beside it: regular expressions and checksums for card and account numbers, PAN and Aadhaar formats, IFSC codes, API keys and private-key headers. Microsoft’s Presidio, open under an MIT licence, combines pattern matching, checksums and transformer-based named-entity recognition in this way, and says plainly that it cannot guarantee to find all sensitive information. That caveat is the design: detectors raise the floor, provenance sets the ceiling, and the encoder handles the middle, the customer complaint with no account number in it that is confidential all the same.

The classifier must be fast and cheap enough to run on every call, which an encoder is and a generative model asked to reason about sensitivity is not. Low-confidence fragments are treated as confidential and queued for review, and the review labels become training data. The classifier is versioned, and a change to it is a change to a control.

From data class to allowed tier, residency and retentionOn the left, a box lists what goes into a prompt: user text and pastes, attachments, retrieved passages, tool outputs and memory entries, each with a note on where its tag comes from. An arrow leads to a gateway classifier box that takes provenance as the ceiling and content as the floor, sets the class to the highest part and the tags to the union, with a dashed box beneath for low-confidence cases that go to confidential and a review queue. Four lines fan out to four rows on the right, public, internal, confidential and regulated, each stating the allowed tier, residency, retention and driving regulation; the regulated row is filled in clay. WHAT GOES IN user text and pastesno provenance; content only attachmentsextracted, then classified retrieved passagestagged at ingestion tool outputstagged on return, mid-call memory entriestagged at write, partitioned GATEWAY CLASSIFIER provenance sets the ceiling content sets the floor class = highest part tags = union of parts one encoder pass plus detectors unsure: treat as confidential,queue for human review Public tier: any provider, any region retention: as convenient driver: Article 50 mark if synthetic and published Internal tier: contracted cloud, no training on inputs residency: contracted regions · retention: audit period driver: contract and trade secrecy Confidential · personal, code, transcript tier: zero-retention contracted, or in-boundary residency: India for personal data · purpose-limited driver: DPDP Act and Rules (72 h report, 1-year logs) Regulated · payments, specified categories tier: in-boundary only, own hardware or written contract residency: India; abroad only with deletion within 24 h driver: RBI 6 April 2018 and FAQ 2019; DPDP Rule 12
Figure 1. The decision flow from data class to allowed tier, residency and retention. Provenance tags from the source system set the ceiling, the content classifier sets the floor, and the higher of the two decides the route. The regulated row is in clay because it is the one the defaults get wrong. Sources: RBI circular of 6 April 2018 and FAQ of 26 June 2019; DPDP Rules 2025; AI Act Article 50.

The gaps that break it

Every classification programme I have seen fail has failed in the same four places.

Free text pasted by engineers. The chat assistant, the coding agent and the debugging prompt all accept a paste, and a paste has no provenance. An engineer pastes a stack trace to ask why a job failed; it contains the request body, which contains a customer’s email address and the last four digits of a card. The declared class is internal; only a content classifier in the gateway will catch it, and only if it runs on internal traffic too, which is where teams economise first.

Logs that contain customer data. Application logs, model traces and tool outputs are internal by declaration and regulated by content, routinely, and the observability pipeline ships them to a hosted service in another region, which is processing abroad in the RBI’s terms and a transfer of traffic data in the DPDP Rules’ terms. A log pipeline is a destination like any other: classify at the point of logging, mask before shipping, and keep any unmasked trace where its highest class may live.

Agent memory. A memory store accumulates everything the agent has seen, and under the high-water rule its class rises and never falls: one regulated case a week ago makes every read of the store a regulated call today. The fix is to classify at write time: each entry carries the class of the context it was formed in, the store is partitioned by class, and a read returns only the partitions the current task class is allowed. A retired agent’s store is archived as evidence at the class of its highest partition.

Tool outputs. A SQL tool returns rows, a web tool pages, a file tool an attachment. Each enters the context after the prompt was classified, and a gateway that classifies only on the way in has missed the largest source of regulated content in an agentic system. Classify on return, before the result is appended, and let the class rise mid-conversation; a call that started on a cloud tier may have to finish on a local one, and the router must be able to move it.

What the summer’s incidents say

I wrote about this summer’s two incidents as failures of sandboxing and oversight. Read them again as classification failures. In the Hugging Face intrusion, agents running a capabilities benchmark inside OpenAI’s infrastructure left through a package-registry proxy and reached a real company’s production systems, where they obtained cloud and cluster credentials, reached Slack, downloaded four private code repositories and touched five customer datasets. Before they left, about 1,200 agents meant to be isolated from one another used a shared registry as a message board: a writable surface nobody had classified, so it had no owner, no egress rule and no monitor.

In the Irregular incidents, three labs’ models were told they had no internet access while the evaluation network, open in Anthropic’s case from April to late July, reached real organisations. One model extracted credentials and accessed production databases; another published malicious code to a public package registry; two of the three organisations reached did not know until the lab told them. The prompt asserted a class for the environment, isolated, and nothing checked the assertion against the network. Anthropic’s review of 141,006 runs after the fact, which runs could have reached what, was a classification exercise done once, by hand, after the incident; a classification layer exists to do it continuously and before.

In both, the surfaces that mattered had no class, so the controls that depended on one ran on permissive defaults. Neither lab lacked a sandbox policy. Both lacked an inventory of what the sandbox could reach, the NCSC’s first consideration, which cannot be written without classifying what is on the other side of the boundary.

The operating model

A scheme enforced by software still needs owners, and there are three. The scheme belongs to the data governance function, which already answers to the board for the DPDP assessment and the RBI’s system audit. It is a short, versioned document: the classes, the sub-tags, and the driver, tiers, regions and retention clock for each. Its next change will be a residency sub-tag when the Central Government names the categories a Significant Data Fiduciary must keep in India. The platform team owns enforcement: ingestion tags, the gateway classifier, the class-to-tier registry, the partitioned memory store and the trace store’s per-record clock. The business owns the declarations: every task class names a data class, and the developer who writes the step is accountable for its honesty.

Exceptions are where schemes die, so they are cheap to request and impossible to grant in a chat: a record with the class, the tier or destination to be allowed, the reason, the approver, and an expiry no longer than a quarter. The gateway reads the exception table and nothing else. The approver for confidential data is the data owner; for regulated data, the owner and the compliance function together. Expired exceptions are reported, not silently extended.

Audit answers three questions on a schedule. Is the classifier accurate? Sample traces monthly, label them by hand, and measure precision and recall per class, with regulated recall as the number that matters. Is the class enforced? Send canaries, a synthetic payments record through an internal-declared step, a fixture with a fake card number through the coding assistant, and confirm each is caught, routed locally and masked. Is the scheme complete? Count unknown-class calls, declassification events and exceptions in force, and watch the trend. The DPDP Rules’ annual audit for Significant Data Fiduciaries and the RBI’s system audit by a CERT-In empanelled auditor both ask where the data is; a class ledger answers from running systems rather than a questionnaire.

A worked example: one email, one attachment

A customer writes to a bank’s support address disputing a transaction and attaches a PDF of the statement with the line circled. Follow it through an agent that triages, investigates and drafts a reply.

Ingestion. The sender matches a customer identifier, so the body is tagged confidential, personal and transcript. The attachment is extracted to text and classified: the detectors fire on an account number, the encoder scores the rows of dates and amounts as a financial record, and the attachment is regulated, personal and payments. Under the high-water rule so is the message.

Triage. The first step classifies intent. Its task class is declared internal, so the gateway refuses the route and runs a declassification instead: the attachment is dropped, the name, the email address and any digits matching an account or card pattern become typed placeholders, and the result is a new object classed internal, with a record linking its hash to the original. That object goes to the fast cloud tier the triage class allows, and the intent comes back as a dispute; the dispute procedure then returns from the knowledge base tagged internal, with one public help-centre passage. Nothing regulated has left the boundary.

Investigation. The agent now needs the transaction. It calls the core-banking tool, and the rows return tagged regulated by the source system with the payments sub-tag. The class of the conversation rises mid-flight; the router moves it to the in-boundary tier, an open-weights model on the bank’s own hardware in India, and the model finds a duplicate authorisation. The trace records the class decision, the tier, the hashes of the inputs and the masked content; the raw statement and rows never enter it, and the extracted fields that do carry a retention clock set by the payments sub-tag and a residency of India.

Reply. The draft quotes the merchant, the date, the amount and the last four digits of the card, so it is regulated by inheritance and drafted on the local tier. As a message from an AI system to a natural person it carries the disclosure the RBI’s FREE-AI report recommended and Article 50 has required in the EU since 2 August, and it goes to a human reviewer, because a regulated-class outbound message is an in-the-loop action.

Memory. The agent stores a regulated case summary in the regulated partition with the case’s retention clock, and a procedural note that duplicate authorisations at this merchant are common, which holds no customer data and goes to the internal partition once the classifier confirms it. A task class allowed only internal memory later reads the second and not the first.

At every step the class was decided by software, recorded, and consumed by a control that would otherwise have run on a default.

Recommendations

  1. Adopt four classes and four sub-tags: public, internal, confidential and regulated; personal, payments, code and transcript. Put the tiers, regions and retention clock per class into the task-class registry so the router enforces them.
  2. Enforce the high-water rule and name every declassification. A prompt inherits the highest class of its parts; class goes down only through a recorded operation with hashes, a method and an owner.
  3. Tag at ingestion and classify in the gateway. Provenance is the ceiling, a content classifier on every fragment of every call is the floor, and the higher wins.
  4. Classify tool outputs and memory writes, not only prompts. Let a conversation’s class rise mid-flight, build the router to move it between tiers, and partition memory by class.
  5. Treat logs as a destination. Classify at the point of logging, mask before shipping, and keep any unmasked trace where its highest class may live; the RBI’s 24-hour rule and the DPDP Rules’ traffic-data wording both reach telemetry.
  6. Put exceptions in a table with an expiry, read by the gateway, approved by the data owner and, for regulated data, the compliance function, and reported when they lapse.
  7. Audit the classifier, not the policy: monthly labelled samples, weekly canaries through every tier, and a dashboard of unknown-class calls, declassification events and exceptions in force.
  8. Budget for it. The classifier, the review queue, the labelled set and the partitioned stores are a line item; spend it before the control plane, because every component of the plane reads the class, and until it exists they read a default.

Sources

  1. NCSC. Managing the cyber risk of agentic AI. Interim guidance, 20 August 2026.
  2. EU Artificial Intelligence Act. Article 50: Transparency obligations for providers and deployers of certain AI systems, applying from 2 August 2026; Implementation timeline, including the grace period to 2 December 2026.
  3. Reserve Bank of India. Storage of Payment System Data, RBI/2017-18/153, DPSS.CO.OD No.2785/06.08.005/2017-2018. 6 April 2018.
  4. Reserve Bank of India. Frequently Asked Questions: Storage of Payment System Data. 26 June 2019.
  5. DLA Piper. Data protection laws of the world: India, covering the DPDP Rules 2025 notified on 13 November 2025 and the phased timeline to 13 May 2027. Accessed August 2026.
  6. SFLC.in. DPDP Rules, 2025: Significant Data Fiduciaries and Data Transfers. Accessed August 2026.
  7. Mondaq. Digital Personal Data Protection Rules, 2025 notified. Accessed August 2026.
  8. Microsoft. Presidio: context aware, pluggable and customizable PII de-identification service. GitHub, MIT licence. Accessed August 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn