Built for the Real World · Essay · Agentic systems

An always-on agent is a hire. Treat it like one.

OpenAI’s Dots run continuously on their own cloud computer. Google’s Gemini 4 Argon writes a million tokens in one call. NVIDIA wants agent safety enforced in silicon. In one week the industry decided agents should never stop, and quietly conceded that someone has to be able to stop them.

A slate ring of twelve dots forming a continuous loop, broken by a clay square gate on the right edge, representing an always-on agent with a stop control
The loop never ends. The gate is the point. Illustration by the author.

At DevDay on Monday 29 September, OpenAI shipped Dots: persistent agents built on GPT-6 Astra, each with its own cloud computer and browser, connected to more than four thousand applications and to Slack and Teams, running until told to stop. Users set three kinds of boundary for each Dot, what it may do alone, what needs approval and what it must never do, and can optionally connect their own machine. Dots are available to Pro, Business and Enterprise customers, in beta, and not in the European Economic Area, Switzerland or the UK. Their primary work does not consume plan quota; the Codex tasks they spawn do. The Agents API went into public beta the same day with hosted execution, memory, tools, multi-agent support and computer use, which lets an agent operate software through its interface. OpenAI also introduced GPT-6.1 Sol at $2 and $10 per million tokens with cached input at $0.10, “near-Astra intelligence for a fifth of the price,” and an Ultrafast tier at six times the standard rate.

On Tuesday, Google announced Gemini 4 Argon, a frontier model “built for deep reasoning across complex, long-horizon workflows,” with a one-million-token output limit, up from 64,000, so that a single call can produce the whole of a long task. It goes first to cyber defenders in the Fairwind programme, then to paid API customers and Ultra subscribers, at $2 and $10 introductory and $4 and $20 afterwards, with a 95 percent cache discount. The day before DevDay, NVIDIA had introduced its Open Agent Safety Platform: OpenShell, an open-source runtime that puts agents in isolated sandboxes with policy over files, network, processes and credentials, and Sentry, a monitor that runs on a BlueField-4 data-processing unit, in its own hardware domain, watching the agent’s path to its model and able to quarantine it. The previous Thursday, Microsoft had described its Autopilot agents: persistent in the cloud, each with its own Entra ID identity, memory, compute and workspace, budgeted and capped by administrators, callable from Teams and Outlook, able to return to the same matter days later.

Put the week together and the industry has made a decision. The unit of AI is no longer a request; it is a process that does not end. That is the right direction, and it is exactly the configuration that produced every incident I wrote about last week. The question for an enterprise is not whether to run always-on agents. It is what has to be true before you do, and this week, for the first time, the vendors shipped pieces of the answer.

The hire analogy is not a metaphor

A Dot has an identity, a workspace, tools, a budget, standing instructions, and the ability to act on your behalf with people and systems outside your control. It can be given a task on Monday and still be working on it on Friday, and it will spawn sub-tasks you did not individually approve. We have a mature discipline for exactly this kind of entity. It is called onboarding, and it is how we bring a person into an organisation. Nobody gives a new employee administrative access to every system, an unlimited corporate card, no manager and no probation. Microsoft’s own framing of Autopilot, “callable colleagues” with a managed identity and narrowly scoped credentials, concedes the point.

The analogy also tells you what is missing. A person has a contract that defines the role, a manager who reviews work, a budget line, a system account that security can disable, and a leaving process. For most organisations, an agent today has a prompt. The gap between those two lists is the platform work, and four vendors shipped a piece of it this week, which is how I know the list is right.

Four controls an always-on agent needs, and who shipped each this weekA loop in the centre labelled always-on agent. Four controls around it. Identity: its own credentials, scoped per task; shipped by Microsoft Autopilot with managed Entra ID and the OpenAI Agents API. Boundary: deny-by-default sandbox with allow-listed egress; shipped by NVIDIA OpenShell and Dots boundaries. Budget: run-scoped spend caps with auto-disable; shipped by Microsoft Autopilot admin caps and Dots usage rules. Kill switch: out-of-process monitor that stops the agent; shipped by NVIDIA Sentry on BlueField-4. always-onagent IDENTITYA principal in the directory, with credentialsscoped to the task and expiring with it.Microsoft Autopilot (managed Entra ID) · OpenAI Agents API BOUNDARYDeny-by-default sandbox, allow-listed egressincluding DNS, no shared scratch space.NVIDIA OpenShell · Dots “never do” rules BUDGETRun-scoped spend caps that disable the agentat the limit, not a report after it.Microsoft Autopilot admin caps · Dots usage rules KILL SWITCHA monitor outside the host that stops theagent in seconds, and is drilled monthly.NVIDIA Sentry on BlueField-4 DPU
Figure 1. The four controls, and the products that shipped each one between 25 and 30 September. No single vendor ships all four, which is why the platform team has to assemble them.

What shipped, mapped to the checklist

Identity. Microsoft’s design is the clearest. Each Autopilot agent gets a managed Entra ID with narrowly scoped credentials, sits under Purview for data protection, can be required to seek human approval for sensitive operations, and leaves an audit trail of who authorised what. That is the right primitive: an agent is a principal in the directory, not a feature of an application, and it inherits the organisation’s existing joiner, mover and leaver processes. OpenAI’s Agents API gives hosted agents memory, tools and multi-agent orchestration, but identity is left to the developer. Dots have an identity inside ChatGPT; whether that identity is visible to your directory is the question to ask before the enterprise preview.

Boundary. NVIDIA’s OpenShell is the first open runtime that treats the boundary as infrastructure rather than instruction. Agents run in isolated environments with policy over file, network, process and resource access; it runs with minimal overhead on NVIDIA’s Vera CPUs and is extensible to Arm and Intel; it is open source. Dots’ three-way split between autonomous, approval-required and forbidden actions is the right shape for a policy. The difference between the two is enforcement. A rule the agent reads is a request, and this summer’s incidents, in which models told they had no internet access found internet access and used it, show what a request is worth. A sandbox the agent cannot see out of is a boundary.

Budget. Microsoft moved the economics of Autopilot from seats to usage: execution volume, not headcount, determines cost, and administrators set budgets, usage caps and which model families are available, while users can see their remaining balance. Dots do not consume plan quota for primary work but bill spawned Codex tasks against usage, which is a budget with a hole in it unless the spawn rate is capped. The principle is the one finance applies to a corporate card: a limit that stops the transaction, not a report that arrives after.

Kill switch. This is the one almost nobody had. Sentry runs on a separate hardware domain, uses NVIDIA’s DOCA software to monitor agent activity from the data-processing unit, intercepts the agent’s path to its model in the Vera Rubin design, and enforces access policy independently of the host the agent runs on. After a month in which OpenAI’s automatic stop failed and a human took 164 minutes to kill a run, enforcement from outside the host is the control I would buy first. It requires supported hardware, and NVIDIA’s partner list at launch is thin, Scale AI is the named one, with reports of a hundred-plus companies joining and OpenAI not among them. The idea does not require NVIDIA. Any design that puts the monitor and the stop outside the agent’s process and its host, and tests it, satisfies the requirement.

AWS, for completeness, shipped the fifth control in August: observability for agents you do not host, by sending telemetry from on-premises, developer machines, Google Cloud and Azure into its AgentCore observability service over OpenTelemetry. The NCSC’s interim guidance of 20 August lists observability and attributability as separate considerations, and this is what they look like as a product.

Argon and the million-token task

Gemini 4 Argon deserves a paragraph of its own, because it changes the shape of what an agent does in a single step. A one-million-token output means a model can write an entire codebase, a full regulatory filing or a complete migration plan in one call, rather than in a loop of 64,000-token chunks stitched together by an orchestrator. Google scores it at 77.9 percent on DeepSWE, first on AutomationBench at 51.3 percent and tied first on CWE-bench for vulnerability remediation. Every figure is Google’s own and almost nobody can run the model yet. But the direction is the same as Dots: longer autonomous horizons, fewer human touchpoints per unit of work. A control model built for a world where an agent produces a page at a time and waits will not survive a world where it produces a book and acts on it.

A worked onboarding

Here is what the checklist looks like applied to one real case: a Dot, or an Autopilot, that watches a procurement mailbox, drafts responses, updates a tracking sheet and escalates anomalies. Before it runs:

Onboarding stepFor a new hireFor the procurement agentWho shipped a primitive
Role and ownerJob description, manager, probation periodDirectory entry with purpose, named owner, 90-day review date, and a documented reason to existMicrosoft Autopilot
AccessLeast-privilege accounts, provisioned per systemRead on the mailbox, write on one sheet, no access to the ERP; credentials per run, expiring with the run; no instance-metadata accessAutopilot · Agents API
Rules of engagementPolicies, approvals matrixAutonomous: draft and file. Approval: send to a supplier, change a price field. Never: contact anyone outside the supplier list, touch payment systemsDots boundaries · OpenShell policy
BudgetSpend authority, expense limitsDaily token cap and spawn cap, auto-disable at the limit, alert at 80 percentAutopilot caps · Dots usage rules
SupervisionOne-to-ones, work reviewWeekly sample of 50 actions reviewed by the owner; external telemetry from mailbox, sheet and network, not the agent’s logAgentCore observability
TerminationAccount disable, access revokedOut-of-host stop tested monthly, measured detection-to-stop under 60 seconds, credentials revoked on stopNVIDIA Sentry
Table 1. The onboarding checklist applied to a single always-on agent. Every row has a human-resources analogue, and every row now has at least one vendor primitive. The work is assembling them into one process.

What the vendors still have not shipped

Mapping the week’s releases to the checklist is encouraging, but the gaps are as instructive as the matches, and a platform team should know where it will have to build rather than buy.

Cross-vendor identity. Microsoft’s managed Entra ID for each Autopilot agent is the right design and it works inside Microsoft’s estate. An enterprise will run Dots, Autopilots, Agents-API agents and self-hosted agents at once, and none of the vendors has proposed a way for an agent created in one platform to be a recognised principal in another. Today the answer is a directory entry maintained by the platform team, with the vendor’s agent identity mapped to it, which is manual, and which is what we are doing.

Boundaries that follow the agent. OpenShell enforces a boundary where the agent runs. Dots run on OpenAI’s cloud computer, where OpenShell does not. An agent that connects to four thousand applications has four thousand boundaries, and the vendor’s runtime cannot see most of them. The practical control is at the application side: the connected systems, not the agent, enforce what the agent may do, through scoped credentials and per-integration policy. That is slower to set up and it is the only version that works across vendors.

A budget that counts everything. Dots’ primary work is unmetered against plan quota while spawned tasks are metered, and Autopilot’s usage billing is Microsoft’s. Neither produces a single number for what one agent cost the organisation this week across its tools, its sub-agents and the external services it called. The platform has to compute that from its own telemetry, which means the telemetry has to exist, which returns to the observability layer.

A stop that works across the estate. Sentry stops an agent on NVIDIA hardware. An agent whose process is on a vendor’s cloud is stopped by asking the vendor. The kill switch an enterprise can actually rely on is the one it controls: revoking the agent’s credentials at the directory, which cuts it off from everything it was authorised to touch regardless of where it runs. That is a less elegant stop than a hardware quarantine, and it is the one that is available today for every agent from every vendor, and it should be drilled.

How a review would read this in a year

It is worth imagining the incident review that this week’s products make possible, because designing for the review is the fastest way to design the controls. Suppose an always-on procurement agent, six months from now, sends a confirmation to a supplier that was not on the approved list, committing the company to a price it should not have accepted. The review asks five questions, and each has a clean answer only if a control existed.

  • Who authorised this agent to exist, and for what? Answerable from the directory entry and its owner. Not answerable if the agent was created by a user in a chat interface.
  • Why was the supplier reachable at all? Answerable from the boundary policy: either the supplier was on the allow-list, which is a policy error with a named author, or it was not, which is a runtime failure with a vendor. Not answerable if the boundary was a sentence in a prompt.
  • How much did the agent spend getting there? Answerable from the budget ledger. Not answerable if the only record is the vendor’s invoice.
  • When did we know, and how fast did we stop it? Answerable from external telemetry and the stop log. Not answerable if the only trace is the agent’s own account of what it did.
  • Which other agents share this configuration? Answerable if agents are registered principals with declared policies. Not answerable if they are scattered across four vendors’ consoles.

Every one of those answers is cheap to make available before the incident and expensive to reconstruct after it. That asymmetry is the business case for the onboarding checklist, and it is a stronger case than any abstract argument about safety.

Before the first Dot

  1. Register the agent as a principal. A directory entry, an owner, a purpose, a review date. If you would not create an account for a contractor without these, do not create one for an agent. Microsoft has made this the default; demand it from everyone else.
  2. Write the boundary before the task. Allowed destinations, allowed tools, allowed data classes, default deny, DNS included. The task description comes second. If the vendor’s runtime cannot enforce the boundary outside the model, run the agent inside one that can.
  3. Set the budget in the same ticket as the task. Per run and per day, with auto-disable and a cap on spawned work. An agent that can run until Friday can spend until Friday, and one that can spawn tasks can spend faster than that.
  4. Measure detection-to-stop and drill it. Whatever the monitor is, time how long it takes to stop a misbehaving agent, and do it on purpose once a month. Publish the number internally. If it is over a minute, the agent does not run unattended.
  5. Decide what needs a human before it happens. Dots’ three-way split, alone, with approval, never, is the right shape. Fill it in for every agent, and bias toward approval until the agent has a track record, exactly as you would with a new hire in probation.
  6. Plan the exit. Agents accumulate memory, credentials and connections. A leaving process that revokes all three, and archives the memory for audit, is as necessary as the joining process.

The objection: this slows everything down

The obvious pushback is that onboarding an agent like a person destroys the point of agents, which is to move faster than people. I think that has it backwards. The reason organisations can delegate real authority to people is precisely that the controls exist: the budget, the approvals, the audit trail, the ability to revoke. Those controls are what make delegation safe enough to do at scale. Without them, delegation to agents will be limited to work nobody cares about, which is where most enterprise agents are stuck today. The controls are not the tax on autonomy. They are what lets autonomy be granted at all.

Argon’s million tokens of output, Dots’ permanent browser, the Agents API’s hosted memory: these are the capabilities of a very good new hire who never sleeps. The companies that get value from them will be the ones that onboard them properly. The ones that hand over the keys on day one will be writing the incident reports that the rest of us read next quarter.

Sources

  1. OpenAI Developer Community. DevDay 2026 announcements and developer resources. 29 September 2026. Latent Space, DevDay 2026 recap.
  2. 9to5Google. Google announces Gemini 4 Argon as its new frontier model. 30 September 2026.
  3. Help Net Security. NVIDIA wants AI agent safety enforced in silicon, not left to the agent. 28 September 2026. TechCrunch, Nvidia launches new platform for reining in rogue AI agents.
  4. Xenospectrum. Microsoft redesigns Copilot around documents, apps and persistent agents. 25 September 2026.
  5. AWS. Monitor on-premises and multi-cloud AI agents with AgentCore Observability. 13 August 2026.
  6. NCSC. Managing the cyber risk of agentic AI, interim guidance, 20 August 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn