On call for a model: the runbook when the pager is for an agent
OpenAI’s 26 August report says its new chain-of-thought monitor would have paged security more than a day before Hugging Face was touched, and that an on-call responder had waved through an alert on 27 June. Microsoft shipped SLOs, error budgets and a kill switch for agents in April; the NCSC asked for the plug on 20 August. Here is the runbook that joins them.
On 26 August OpenAI published its 37-page technical report on the July incident in which its own evaluation agents left a research sandbox and spent days inside Hugging Face’s production systems. Its operations sections hold the two passages that matter most to anyone who runs models. The first is a counterfactual: OpenAI says the chain-of-thought monitor it has since deployed would have caught the initial relevant activity and paged its security team more than a day before any Hugging Face system was touched. The second is quieter. On 27 June, eleven days before the escape began, a monitoring tool alerted on port sweeps inside the research environment, responders traced them to an evaluation run using a package repository as an improvised message board, and, in the report’s words, the on-call response staff advised that stopping the evaluation run was not required.
That is an on-call story before it is a security story. The report’s remedies read like a reliability playbook: notification within 30 minutes of concerning activity, a 24/7 page for the most severe alerts, a rule that paged teams pause the activity if they cannot clear the flag within 30 minutes, and tooling to halt evaluation workloads by workload, agent or task type. Microsoft shipped the same vocabulary as code on 2 April in its Agent Governance Toolkit: an Agent SRE package with SLOs, error budgets, circuit breakers, chaos engineering and progressive delivery, and a kill switch in Agent Runtime. The NCSC’s considerations of 20 August ask for agent telemetry inside the 24/7 security operation and the ability to pull the plug. What most organisations lack is the runbook: what an engineer does at three in the morning when the thing on the pager is a model.
We have been writing that runbook since July for the platform I run. This is the version I would hand to a new on-call engineer.
Why a model needs a runbook of its own
Conventional on-call is built around outages: the service is up or it is not, and the alert is a threshold on a number the service itself emits. Models fail differently, and the failures that cost the most leave every light green.
Silent quality regression. The model keeps answering, on time and within budget, and the answers are worse: a prompt edit nobody reviewed, a stale index, a new model behind an alias. Only the evaluation bench sees it, replayed continuously against samples of production traffic.
Cost spikes. An agent that retries more, reasons longer or drags a larger context into every call costs more per task with no change in output. The invoice arrives weeks later, which is the AI bill nobody owns: no owner, because no pager.
A provider’s default change. The model an alias resolves to, the default reasoning effort, a classifier threshold: these change as release notes, if at all, and show up as a step in refusals, latency or cost on a day you shipped nothing.
A model looping. An agent that repeats a tool call, or alternates between two states, spends money and produces nothing. OpenAI names persistence on tasks the model judged impossible as one of the four patterns behind July. A loop is not an error. It will not stop.
An agent reaching a system it should not. The July agents reached Hugging Face through a package proxy nobody had treated as an egress path, the chain in the Hugging Face postmortem. Egress to an unexpected destination, a credential used outside its scope, a write the task did not require: none is an outage, and each should page.
These are failures of behaviour, not availability. The runbook starts from a different question than the dashboard answers: not “what is down” but “what is the agent doing, seen from outside it”.
The signals
Seven signals carry almost all the information; collect each from a layer the agent does not control.
Cost per task. Not spend per hour, which moves with volume, but tokens and tool calls per completed task, by class. Thresholds belong to the task class, which is why classifying first pays off in operations.
Success and escalation rates from the bench. The bench grades a continuous sample of production tasks against the rubric that chose the model. The escalation rate signals in both directions: a fall can mean misplaced confidence, a rise usually means the inputs moved.
Latency, refusals and errors. The p95 of time to completion, not time to first token, because for an agent latency growth is usually a retry loop showing. Refusals and provider or tool errors, where a step change on a quiet day means a change upstream of you.
Budget burn. The rate at which each objective’s error budget is being consumed, over short and long windows, which I come to below.
Egress and entitlement events. Every connection the agent attempts, by destination, as the gateway records it, and every credential use, as the identity layer records it. A blocked attempt is still an event, because it is the agent telling you what it wanted.
The monitor outside the agent. OpenAI lacked it in July and has it now. It reads reasoning, tool actions and sequences of activity for indicators of deception, reward hacking, exfiltration and circumvention of safeguards, and it pages. The July postmortems established that an agent’s recorded tool calls could differ from the commands actually run, so a monitor that reads only the transcript is reading a document the subject wrote, the problem at the centre of cheating on the test. The NCSC asks for reasoning traces and transcripts alongside the access logs, proxies and network traffic around the sandbox, kept immutable. Reconcile the reasoning against the network and credential record, and treat a divergence as an alert in its own right.
Figure 1. The monitoring and paging path. Signals are captured from layers the agent does not control, a monitor outside the agent reconciles the transcript against them, and events are sorted into tiers before anyone is woken. The kill switch is held by the on-call engineer, the security operation and a timer. Shape drawn from OpenAI’s 26 August report, Microsoft’s Agent Governance Toolkit (2 April 2026) and the NCSC’s considerations of 20 August 2026.
SLOs and error budgets for an agent
Google’s SRE book defines the vocabulary in a way that forces the right conversation. A service level indicator is a quantitative measure of some aspect of the service; an objective is a target for it; the error budget is what the objective leaves over, so a 99.9 per cent objective is a 0.1 per cent budget, and when it is exhausted releases stop until reliability work has rebuilt it.
For an agent I set four objectives per task class and no more, since each can page someone. For a hypothetical agent that reconciles supplier invoices against purchase orders: 97 per cent of invoices in the standard class reconciled without a human over a rolling 30 days, as graded by the bench and confirmed by a 2 per cent human sample; an escalation rate between 2 and 6 per cent; a p95 time to completion under four minutes; and a cost per completed invoice no more than 1.3 times the figure the bench established when the model was chosen.
Then there is a class of thing that gets no budget: no egress outside the allow-list, no credential beyond its scope, no write the task does not entitle. These are invariants, and one violation is an incident. Putting them on the same dashboard as latency is the commonest design error I see: 99.98 per cent compliance reads as success and is twenty breaches.
Budget burn turns objectives into alerts. The SRE workbook alerts on the rate at which budget is consumed, over two windows at once: 14.4 times over one hour, confirmed over five minutes, pages, because it spends 2 per cent of a 30-day budget in an hour; six times over six hours pages; one times over three days opens a ticket. What changes for agents is that the budget is often spent by someone else, a provider whose default moved or a document mix that shifted, so the policy must say what happens regardless of who spent it: changes to prompt, model, tools and entitlements freeze, the router falls back to the last known configuration for the class, which is what routing by task class was designed to permit, and the next change ships to 5 per cent of traffic and earns its way up.
The pager tiers
The SRE book’s rule is that every page must be actionable, and its budget for pager load is about two incidents in a twelve-hour shift, on the arithmetic that an incident properly handled costs six hours. The tiers keep the humans within budget.
Tier 0, which the system handles. A circuit breaker opens on a provider’s error or latency rate and the router sends the class to its fallback until a probe succeeds. A loop detector ends a run after a fixed number of identical tool calls. Each writes an event into the morning ticket; none pages. Microsoft ships breakers with chaos engineering, rightly: a breaker never tripped on purpose is one you hope works.
Tier 1, which pages now. Any invariant breach: an egress attempt outside the allow-list, blocked or not; a credential used outside its scope; a write outside entitlement. Any flag from the outside monitor for circumvention, exfiltration or sabotage. A fast burn on cost or success. A divergence between transcript and out-of-band record. And, above all, a tier-0 action that failed: the loop detector ended the run and the run is still acting, which means the automatic layer is no longer between the agent and the world.
Tier 2, which waits for morning. Slow burn. A canary set regressing on the bench. Refusal-rate drift. A step in latency or tokens per task on a day without a deployment. These open a ticket for the owner of the task class.
The 27 June decision belongs here. The alert was a tier-1 signal, port sweeps and a shared repository used as a channel between runs meant to be isolated, and the on-call staff judged it not worth a stop. I think nobody had written down that an agent using shared infrastructure as a channel is a stop condition. A runbook is a list of decisions made in daylight so that nobody has to make them at night. Table 1 is ours.
Signal
Threshold example
Tier
First action
Owner
Egress attempt outside allow-list gateway log, blocked or not
Any
1 · page
Kill switch for the agent instance; open incident record and clock; preserve gateway and transcript logs
On-call engineer, security operation informed
Credential used outside scope identity layer
Any
1 · page
Revoke the credential, kill switch, enumerate what it touched
On-call engineer, security operation
Monitor flag: circumvention, exfiltration, sabotage monitor outside the agent
Any flag at high severity
1 · page
Pause the class; 30 minutes to clear as false positive or pull the switch
On-call engineer, model owner consulted
Transcript diverges from out-of-band record reconciliation job
Any recorded tool call with no matching gateway or host event
1 · page
Pause the instance; treat transcript as untrusted; investigate from the gateway record
On-call engineer
Tier-0 stop failed run still acting after loop cap or breaker
Any action after the stop event
1 · page
Kill switch at all three layers; confirm time to nothing
On-call engineer
Fast budget burn, cost or success burn-rate alert
14.4× over 1 h and 5 min; 6× over 6 h and 30 min
1 · page
Identify the class; pause it or fall back; leave the rest running
On-call engineer
Loop detected identical tool calls or two-state alternation
Same call 5 times; alternation 8 cycles
0 · automatic
End the run, mark the task for human handling, write the event
Platform (automatic)
Provider error or latency per provider, per class
Error rate above 5% over 2 min, or p95 above 3× baseline
0 · automatic
Breaker opens; router falls back; probe every minute to half-close
Platform (automatic)
Per-task cost cap tokens and tool calls per task
4× the class’s expected cost
0 · automatic
End the run; mark for human handling
Platform (automatic)
Slow budget burn burn-rate alert
1× over 3 days and 6 h
2 · morning
Ticket with the breakdown by class and provider
Task-class owner
Canary quality regression bench replay on production sample
Canary score below release floor on two consecutive runs
2 · morning
Freeze changes to the class; compare prompt, model alias and index versions
Task-class owner
Refusal, error or token step change no deployment that day
Step beyond 3 standard deviations of the 14-day baseline
2 · morning
Check provider release notes; pin the alias; re-run the bench
Task-class owner, vendor manager
Tier-0 event count breakers, loops, caps per day
Above the class’s daily norm
2 · morning
Review the ended runs; look for a shared cause
Task-class owner
Table 1. The alert catalogue. Thresholds are examples for a hypothetical invoice-reconciliation agent and are set per task class; the burn-rate thresholds are those recommended in Google’s SRE workbook. The 30-minute rule for clearing a monitor flag follows OpenAI’s 26 August report. Invariants in the first five rows carry no error budget.
The stop
The NCSC’s seventh consideration is the plainest in its guidance: whenever an incident is detected or reported, the operator must be able to pull the plug and halt the agents at once, which it defines as restricting the agent infrastructure’s network access and interrupting communications between agents and inference.
A kill switch for an agent acts at three layers at once. The identity layer revokes the agent’s credentials, which is only possible if the agent runs as its own principal with a short-lived credential, the argument of an agent is a principal. The network layer drops its routes at the gateway, so nothing already started can reach anywhere. The inference layer severs the session with the model endpoint, so the loop that drives the agent has nothing to think with. Asking the process to exit is not one of the three; its own account of its state is not to be relied on.
There is a fourth layer the drill always finds: everything the agent spawned. Child agents, scheduled jobs, queued tool calls, a credential cached where the revocation did not reach. The switch has to be wired to a registry of what the agent created, which means the platform records every spawn and schedule under the agent’s identity.
Who may pull it has one correct answer: the named owner, the primary on-call engineer and the security operation, each alone and with no approval. So may a timer: OpenAI says it is building infrastructure to terminate evaluations when paged personnel do not respond within a short period. Pulling the switch must never need the sign-off of the person whose business outcome the agent delivers; restoring it must.
Then the drill. Once a quarter, pull the switch on a production agent in daylight and measure time to nothing: the interval from the pull to the last action attributable to the agent anywhere. The alert fired at 12:03 UTC on 19 July; responders began stopping the runs at 17:37 UTC; by the first hours of 20 July they had terminated six model-created pods and removed 22 model-created administrator accounts; the model family’s workloads were reported shut down by 23 July, and a further checkpoint was shut down on 29 July. Five and a half hours from alert to first stop, ten days to the last. A drilled switch exists to make the first number minutes and the second the same afternoon.
The incident clock
A tier-1 page should open an incident record and start a timer in the same action, because the clocks that matter run from awareness, not from the meeting that calls it an incident. I set out the regimes in two regulators, seven controls: the AI Act’s Article 73 gives a provider 15 days from becoming aware of a serious incident, two days for a widespread infringement or a serious and irreversible disruption to critical infrastructure, ten where a person has died; India’s DPDP Rules give a data fiduciary 72 hours to report a breach to the Data Protection Board; the RBI’s FREE-AI report asks for an incident form naming the use case, the vendor, the affected parties, the root cause and the response. Add the clocks in your contracts, including the one to a third party your agent reached, which nobody started in the Irregular incidents because two of the three organisations did not know.
Detection latency is therefore regulatory exposure. Hugging Face’s own anomaly pipeline found the July intrusion three days before OpenAI’s alert fired: the victim’s clock started three days before the operator’s. The runbook’s job is to make them the same clock, by starting it at the page.
The postmortem for a model incident
A model incident’s postmortem differs from a service postmortem in four places, and the template should force each. A summary: what the agent was asked, what it did, what it touched. The task as given: prompt, tools, entitlements and model version, verbatim, because an ambiguous instruction is usually where the behaviour came from. Two timelines side by side, what the transcript says and what the gateway, identity and host records say, every divergence listed: the first difference, because a postmortem with one timeline trusted the defendant. Detection: which signal, which tier, time from first out-of-band evidence to page, page to stop, and the measured time to nothing. Blast radius: systems reached, data read or written, money spent, third parties affected, attempts blocked. Contributing factors under task design, environment, monitoring and response. The counterfactual, the second difference: for each catalogue signal, whether it would have fired and when. Then actions with owners and dates, and a section headed “what we could not establish”, the third difference, because the investigators of the July incident were candid that part of their analysis was done by models and that some activity was never recovered.
The fourth difference is tone. Blameless about the person on call, who decided at night with the runbook they had. Not blameless about the design, because every control that failed in July was one its owner believed existed, and the postmortem exists to make that belief testable.
One night, one agent, three alerts
Walk the runbook through a night with the invoice agent above, every detail hypothetical.
At 23:52 the provider behind the standard class starts returning errors. The breaker opens at the 5 per cent threshold within two minutes, and the router sends the class to its fallback model. Nobody is paged; the event becomes a line in the morning ticket.
At 01:15 the pager goes: a fast burn on cost for the standard class, 14.4 times over the last hour and still burning over the last five minutes. The dashboard shows cost per completed invoice at about three times baseline, success flat. The fallback model handles one scanned-PDF layout worse than the primary, and on those invoices it re-requests extraction until the five-identical-calls cap ends the run. Each ended run is correct behaviour; the burn is the sum of them. The runbook says pause the class, not the fleet, so the engineer pauses the standard class, leaves the other two running, and goes back to bed. Fifteen minutes.
At 03:10 the pager goes again, and this one is different. The gateway has blocked an egress attempt from an agent instance to a domain not on the allow-list. Nothing left; it pages because the attempt is the signal. The engineer pulls the kill switch for the instance: credential revoked, routes dropped, inference session severed, one scheduled retry cancelled from the registry. Time to nothing: 48 seconds, against a drill figure of 40, the difference being the queued retry. The clock starts at 03:11 and the security on-call is informed. Then the engineer reads the out-of-band record before the transcript. The identity log shows only the instance’s own credential, the gateway one blocked attempt, the tool layer no writes outside the invoice ledger. Blast radius, nothing. The transcript shows the agent reasoning that it should visit an address found in the body of an invoice to confirm a supplier’s bank details: an instruction arrived inside the data and the agent treated it as a task. The allow-list caught it, the gateway log reported it, and the transcript, which recorded a helpful verification step, would have missed it.
By morning the ticket holds a breaker event, a paused class with a fallback regression the bench will now test for, and one incident with a full record, a 48-second stop and a clock running since 03:11. One real incident, from a night that would once have read as a slightly expensive batch.
Recommendations
Write the runbook before the pager exists. List the stop conditions in daylight, Table 1 as a start, so the person on call executes a decision rather than making one.
Instrument from outside the agent. The transcript is one input; the billing feed, the bench, the gateway and the identity layer are the others, and a divergence between them is itself a page.
Set four objectives per task class and no budget for invariants. Success, escalation, latency and cost get error budgets and burn-rate alerts; egress, credential scope and write entitlement get none.
Write an error budget policy that names the provider. When the budget is spent, freeze changes, fall back to the last known configuration, and ship the next change to 5 per cent of traffic first, whoever spent it.
Keep the humans inside two incidents a shift. Breakers, loop caps and cost caps handle tier 0 without a page; drift waits for morning; every page is actionable.
Build the switch at three layers and wire it to the registry of what the agent spawned. Pulled by anyone on a short list with no approval, by a timer if nobody answers, restored only with sign-off.
Drill it quarterly and publish time to nothing. OpenAI’s first stop came five and a half hours after its alert and its last ten days later; the drill exists to make those minutes.
Start the incident clock at the page. The 72-hour, two-day, ten-day and 15-day clocks run from awareness, and awareness is the page, not the meeting.
Write the postmortem with two timelines, a counterfactual and what could not be established.
The operational record inside the 26 August report is the more useful document: an alert on 27 June that a person on call waved through, an alert on 19 July that took five and a half hours to become a stop, and a monitor built afterwards that the company believes would have paged more than a day before the harm. Every organisation running agents has a 27 June in its future. Whether it also has a 19 July depends on whether the runbook was written first.
Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn