On Wednesday 24 June, Google made computer use a built-in tool of Gemini 3.5 Flash, in public preview through the Gemini API and the Gemini Enterprise Agent Platform. Until then, driving a screen with a Google model meant calling the separate Gemini 2.5 Computer Use model of 7 October 2025, which did not control a desktop. The new tool takes a screenshot in and returns an action out, in browser, mobile and desktop environments, and every action carries an intent field stating why the model chose it. It ships with seven safety policy categories that can allow an action, require confirmation or block it, an optional detector that stops the task when it finds a prompt injection on the screen, and the same model the application already calls for search, maps and function calls. Google’s model card puts Gemini 3.5 Flash at 78.4 on OSWorld-Verified against GPT-5.5’s 78.7; the price list puts it at $1.50 and $9 per million tokens against GPT-5.5’s $5 and $30, and the tool costs nothing extra, since screenshots and actions are charged as ordinary tokens.
As a capability announcement it is modest. Claude Opus 4.8 scored 83.4 on the same benchmark on 28 May, and OpenAI has exposed a computer-use model in its API since 11 March 2025. As a price announcement it is the release of the month. When the cheapest tier of a frontier family can operate a screen at a third of the frontier price, every application without an API becomes automatable by anyone with a browser and a budget. Enterprises have been here before. It was called robotic process automation, and its lesson is about to be relearned at ten times the scale, by agents more robust than the bots they replace and considerably less predictable.
Why built-in matters
A standalone computer-use model makes the screen a separate department: two models that share only whatever text was passed between them. Built-in means the agent that read the policy document, called the pricing function and searched the web is the same agent that now opens the vendor portal, with the whole task in its context.
Three details matter more than the headline. First, the intent field: each returned action comes with the model’s stated reason. Second, safety is configuration rather than prose. The seven categories are financial transactions, sensitive data modification, communication tools, account creation, data modification, user consent management and legal terms, each resolving to allowed, require_confirmation or blocked; a developer can switch categories off with disabled_safety_policies. Prompt-injection detection is a boolean, enable_prompt_injection_detection, and it defaults to false. Third, the vendor’s caveat. Google’s documentation says the preview “may contain errors and security vulnerabilities,” and advises against critical decisions, sensitive data or actions whose errors cannot be corrected. That is this post’s risk model in miniature, written by the people who built the tool.
One clarification. The 78.4 is the score the model card carried when Gemini 3.5 Flash went to general availability on 19 May; the 24 June release published no new number. The tool is new, the ability is five weeks old, and what changed is that it is now in the production model at the production price.
The RPA lesson
Blue Prism named the category in 2012; UiPath, founded in Bucharest in 2005 as DeskOver, shipped its first desktop automation product in 2013. The promise was Wednesday’s: software that uses software the way a person does, so the system of record never changes and the integration is never built. EY’s 2016 paper on getting ready for robots estimated that 30 to 50 percent of initial RPA projects fail. Deloitte’s 2017 survey of 400 organisations found 63 percent had missed delivery deadlines. By 2019 the trade had named the cause: Andy Walter, formerly of Procter and Gamble, wrote in CIO that “breaking bots” were the first enemy of RPA success, broken by interfaces being optimised, data formats evolving and connected systems being upgraded.
The mechanism matters because the new tools change it. A bot was a script anchored to selectors and coordinates: click the third button in the second panel, paste into a field with a known identifier, read the number at a known offset. A renamed button or a slower server and the bot clicked thin air, usually at two in the morning, and a developer was paged. The cost was never the licence. It was maintenance, which scaled with the number of screens touched, which is why most programmes stalled at a few dozen bots.
A vision-driven agent fails differently, and both halves of the difference matter. It is more robust, because it sees the screen roughly as a person does: a moved button is found, a renamed field is read, an unexpected dialog is dismissed. It is also less predictable, in three ways a bot never was. The same screen on two days can produce two sequences of actions. The agent can take a path nobody scripted, because it chooses rather than replays. And it can be talked to by the screen, which a bot could not. A bot did exactly what it was told until the interface changed; an agent does roughly what it was told whether or not the interface changed. The second is more useful and harder to govern, and UiPath is a named launch partner for Google’s tool, which tells you where the RPA industry thinks its category goes next.
Then there is scale. Each bot was a project: a process map, a developer, a licence, a change board. If the marginal automation falls from a six-week project to an afternoon with a prompt, the count rises by an order of magnitude and the review each one gets falls by about as much. That is what I mean by RPA’s risks at ten times the scale: the individual bot is better, the population is larger and harder to enumerate, and the controls RPA programmes eventually built, a registry, a credential vault, a change process, are needed in the first week rather than the fifth year.
Where it earns its keep, and where it is the wrong tool
Legacy systems with no API. Every large enterprise runs core systems whose vendor is gone, or charges for an interface it never finished, or exposes a thick client and nothing else. That is where the retyping happens today, between a screen on the left and a screen on the right, and a desktop-capable agent is the first tool since RPA that reaches them and the first that survives a patch to the client.
Vendor and counterparty portals. Third-party administrators, logistics trackers, government filing sites, bank and supplier portals. Each is a login and a form, the counterparty will never build you an API, and the work is pure transcription.
The long tail of internal tools. The three hundred applications with twelve users each, on a framework nobody maintains. Nobody will add an API to them; a screen-driving agent is the cheapest integration they will ever get.
Anything with an API is the wrong place. Under Google’s documented tiling rule, an image larger than 384 pixels on a side is cut into 768-pixel tiles at 258 tokens each, so a 1,280 by 800 screenshot costs 1,032 tokens before any reasoning, and the loop sends one per step. A function call costs a few hundred tokens and returns in milliseconds; Google quoted around 225 seconds per task for its 2.5 model on the Browserbase harness. A typed interface is deterministic; a screen is interpreted, and 78 percent on a benchmark means one task in five did not finish. My rule has three clauses. If a function exists, call it. If the vendor could ship one in a quarter, ask, and screen-drive meanwhile with a date attached. If neither, computer use is the right tool: in the task-class registry I described on Wednesday, screen-driven is its own class with its own model, step budget, approval policy and recording requirement.
The risk model
Prompt injection through the screen. On 20 August 2025, Brave disclosed that Perplexity’s Comet browser handed page content to its model without distinguishing it from the user’s instructions; the demonstration extracted an email address and a one-time code from a signed-in session. Five days later Anthropic published numbers for Claude for Chrome: across 123 adversarial cases in 29 scenarios, attack success fell from 23.6 to 11.2 percent with mitigations, and a browser-specific challenge set from 35.7 percent to zero. Google now ships a detector that halts the task, off by default. A screen is an input channel, everything on it is untrusted, and the model is the only thing telling the user’s instruction apart from a sentence in a claim note that says to mark the claim as matched. The best published mitigation lets one attack in nine through.
Credentials in the session. The agent works inside a logged-in session, and whatever that session can do, an injected instruction can do. An agent signed in as a claims supervisor can approve a claim. Anthropic’s documentation advises against giving the model credentials at all. So the agent gets its own account with the narrowest role the target supports, and the secrets live in the harness that establishes the session, never in the context the model reads.
Irreversible actions. Submit, pay, delete, send, agree. OpenAI’s Operator asked for confirmation before acting from its first day in January 2025, and ChatGPT agent added a watch mode and blocked financial platforms when it absorbed Operator that July. Anthropic’s guidance names cookies, financial transactions and terms of service as actions needing a human. Google’s categories make the same list configurable. All three have converged on one shape: the model proposes, something else disposes. In an enterprise the something else has to be a policy engine that knows which button on which screen releases money, because nobody is watching two thousand sessions.
Audit of what the agent saw and clicked. In RPA the log was the script. Here the why lives in the model’s reasoning and the what lives in pixels, so the log must be the screenshots, actions, intents and gate decisions, kept together. The intent field helps, and it is the model’s account of itself, which is not evidence. The screenshot is.
The controls
None of these is novel. Each is an RPA control or a browser-security control, moved to where an agent lives.
A separate browser profile with its own identity. Google’s documentation suggests a VM, a container or a dedicated browser profile, and the profile is the minimum: a fresh state with no personal cookies, password manager, extensions or history. The identity is an account the target system knows as the agent, with the narrowest role, rotated and revocable. Never a human’s session: that gives the agent a human’s permissions and an audit trail that blames the human.
Allow-listed domains, enforced in the network. An egress proxy for the sandbox that reaches the target domain and nothing else, with denials logged. Claude for Chrome’s site-level permissions are the consumer version; the enterprise version lives in the proxy, because a rule the model reads is a request and a rule the network enforces is a boundary. When an injected instruction sends the agent to a URL, the proxy refuses, and the refusal is a security event worth paging on.
Approval gates for irreversible actions. Start from the vendor’s categories and add the specific controls on specific screens that release money, submit filings, delete records or send messages. Default to confirmation, route the request to a human queue with the screenshot attached, and never let the approving step run on the model that proposed the action.
Screen recording as the audit log. Keep every screenshot, action, intent and gate decision for the retention period of the process. This is the record a regulator will ask for, the evidence an incident review needs and the dataset for your own benchmark.
Per-task budgets. Steps, tokens, minutes and money, each with a cap that aborts the task into a human queue rather than retrying. An agent at benchmark accuracy fails one task in five, so a failure has to be cheap, and a queue a person works through in the morning is cheap.
Benchmark literacy: what OSWorld-Verified measures
OSWorld is a set of 369 tasks on an Ubuntu virtual machine, across LibreOffice, Chrome, GIMP, VLC, Thunderbird, VS Code and the operating system. Success is judged by execution: 134 evaluation functions inspect the resulting files, settings and application state. The human baseline is 72.36 percent. OSWorld-Verified, published on 28 July 2025, fixed around 300 reported issues, added fuzzy matching so formatting differences do not fail a correct answer, and moved the harness to AWS with fifty-fold parallelism.
Now read Table 1 against what the benchmark excludes. There is no authentication; every task starts signed in. There are no consequences; a wrong click costs nothing. There is no mobile, Windows or macOS, and no portal redesigned last Tuesday; The Next Web’s report on Wednesday’s release notes that these models still struggle with unexpected pop-ups, CAPTCHAs and unfamiliar layouts. A score above 72 says that on a fixed set of desktop tasks the model finishes more than a human does. The vendors do not run it identically either: Anthropic changed its harness and restated Opus 4.7 to 82.3 when it published 83.4 for Opus 4.8, and Google’s 78.4 is a May number quoted in June. Three figures within five points, measured three ways, say the capability exists and nothing about which model to buy.
The answer is the one the intent era got right, as I argued in the history I wrote on 4 June: evaluate on your own traffic. Fifty tasks on your own screens, run against every candidate and re-run on every release. Measure success rate, steps per task, cost per task and how often the policy gate fires. Those four numbers choose the model; the public benchmark only decides who gets on the bench.
A worked example: reconciling claims across a portal
An insurer has a block of health claims administered by a third party. The administrator’s portal shows each claim’s status and paid amount; the insurer’s core system holds what the amount should be; reconciliation means comparing the two for every claim and flagging mismatches. The portal has no API and the administrator has declined, twice, to build one. A team pastes claim numbers into a search box, two thousand a day. Pure transcription, a counterparty that will not integrate, and work that must be correct and recorded rather than fast.
The design follows Figure 1. The core system is reached by a function call, never a screen. The agent runs in a sandboxed browser, signed in to the portal with its own read-only account, behind a proxy that allows the portal’s domain and nothing else. Per claim it searches the number, opens the result, reads the status and paid amount, and writes both to a reconciliation table through a function. The comparison is done in code against the core record, not by the model reading two numbers, which removes the obvious injection target, a claim note telling the agent the amounts match. A step budget of fifteen aborts the claim into a human queue. The financial-transactions and communication categories stay at confirmation, the portal’s resubmit button is added to the gate by name, and the read-only role sits under both.
The cost is arithmetic from the published prices and the tiling rule. Call it eight model steps per claim. Each step sends a cached prefix of about 2,000 tokens, the system prompt and tool definitions, at $0.15 per million; roughly 3,500 fresh input tokens, the current 1,032-token screenshot, two retained earlier ones and the text between, at $1.50; and about 600 output tokens of reasoning, intent and action at $9. That is about 1.1 cents a step, nine cents a claim, $180 a day and under $4,000 a month. The same loop on GPT-5.5 at $5 and $30, with cached input at $0.50, is about 3.7 cents a step and 29 cents a claim, a little over three times; Opus 4.8 at $5 and $25 lands in the same band. Batch pricing, which halves Gemini 3.5 Flash to $0.75 and $4.50, does not apply, because each step depends on the previous screenshot; the way to use it is to push everything that does not need a screen, the comparison and the daily report, out of the loop. The token counts are estimates; the structure is not. Screenshots dominate the input, reasoning dominates the output, and a model that takes fewer steps beats a model with a lower price.
The failure arithmetic matters as much. At benchmark-like accuracy, four hundred claims a day fall into the human queue in the first week, which is the design working; by the fourth week the measured rate on these screens is the number that matters, and it will not be 78.4. The recording is sixteen thousand screenshots a day, kept to the claims retention standard: the evidence for every flagged mismatch and the test set for the next model. The team that pasted claim numbers now works the queue and reviews a sample of matches, which is a better job and still a job, because an agent at this accuracy with this blast radius does not run unattended.
Recommendations
- Inventory the screens first. Sort every process where a person retypes between systems: an API exists, an API could exist within a quarter, no API will ever exist. Only the last column belongs to computer use; the middle column needs a date.
- Make screen-driven a task class. Model, step budget, approval policy, allow-list and recording are properties of the class, not the prompt, and the router chooses the model, so Wednesday’s price is a configuration change.
- One profile and one identity per agent. A fresh browser profile, an account the target knows as the agent, the narrowest role, secrets in the harness, revocation in one step. Never a human’s session.
- Enforce the allow-list in the network. Default deny at the proxy, the target domain only, denials logged and alerted. A rule the model reads is not a control.
- Turn the detectors on and do not rely on them. Google’s injection detection defaults to off. Switch it on, treat every pixel as untrusted, keep comparisons and decisions in code, and assume one attack in nine still gets through.
- Gate the irreversible by button, not only by category. Map the controls that move money, submit, delete or send; require confirmation into a human queue with the screenshot; keep the approver separate from the proposer.
- Keep the recording. Screenshot, action, intent and gate decision for every step, for the process’s retention period. It is the audit log, the incident evidence and the benchmark data.
- Budget per task and run your own benchmark. Caps on steps, tokens, minutes and money that abort into a queue; fifty tasks on your own screens, re-run on every release; success, steps, cost and gate hits published monthly. Plan for ten times the bots RPA ever reached, because that is what a third of the price buys.
For fifteen years the browser was what a person used when the API did not exist. From this week it is the API, for every application in the estate, at a price that makes the integration backlog look optional. The enterprise has also acquired, in one preview release, an interface to every system it owns and every portal it logs into, with an agent on the other side that reads what it is shown and does roughly what it is asked. RPA taught us that bots break. The lesson this time is that they do not, and the controls have to be ready before the second hundred are running.