Built for the Real World · Essay · Model strategy

Four frontier models in three days, and the one change that matters

Anthropic, Google, OpenAI and Meta all shipped between Monday and Wednesday. The headline is the pace. The structural change is that the most capable tiers are now gated by who you are, not what you pay.

Four overlapping rings in clay and slate on a warm paper background, representing four frontier model releases in one week
Four releases, one week. Illustration by the author.

On Monday 1 September, Anthropic released Claude Fable 5.1 and a twin called Mythos 5.1 that shares its weights but not its safeguards. On Tuesday, Google shipped Gemini 3.8 Flash and a variant called 3.8 Flash Cyber that only vetted defenders can call. On Wednesday, OpenAI released GPT-6 Astra, the first model it has ever classified as “Critical” for cyber capability under its own framework, and Meta quietly put Muse Spark 1.3 on the market at a blended price near ten cents per million tokens. Four frontier labs, three days, five models, and at least three of them not fully available to anyone with a credit card.

I run the AI and data platform for a large industrial group. Weeks like this generate the same question from every subsidiary, usually within hours: which one do we switch to? I want to argue that this is the wrong question, and that the week contains one structural change that will outlast every model in it. But first, the facts, because the details are where the decisions live.

What actually shipped

Claude Fable 5.1 is Anthropic’s flagship, priced at $10 per million input tokens and $50 per million output, unchanged from Fable 5. The number that changed is the cache-read rate, cut by three quarters to $0.25 per million, with batch processing at half the standard rates. Anthropic’s own estimate is that typical workloads get about 25 percent cheaper and highly agentic workloads about 45 percent cheaper, entirely from the cache change, because an agent that re-reads a long context on every step pays mostly for cache reads. The model has a one-million-token context window and 128,000 tokens of output, and exposes five effort levels from low to max that trade latency and cost for quality.

The benchmark table is a real step rather than a point release. On Terminal-Bench 4.0, a multi-step agentic coding suite, Fable 5.1 scores 55.8 percent against 42.0 for Fable 5 and 52.3 for Opus 5. On Terminal-Bench-Science it more than doubles its predecessor, 52.6 against 24.7. On GDPval-AA, which scores professional knowledge work against human experts, it posts 1853 against Fable 5’s 1723. On OSWorld 2.0, the computer-use benchmark, 77.9 percent. These are Anthropic’s numbers, but they are consistent with what early customers say: Square reports the model was “far more efficient per token than Opus 5” on a 30-day simulated run-a-business evaluation with full access to simulated tools, customers and vendors.

Two safeguard changes matter for enterprises. Biology refusals on benign medical queries fire 85 percent less often than on Fable 5, and cyber safeguard interventions per session are down roughly 60 percent, with vulnerability identification now permitted for defensive work. Anthropic also introduced what it calls Enterprise Frontier Safeguards: zero-data-retention running on customer-controlled infrastructure, which is the configuration our regulated subsidiaries have been asking for since the first Claude.

Claude Mythos 5.1 is the same model with reduced cyber and biology safeguards, at the same price, available only through two vetting programmes: a Cyber Verification Program for defensive security teams and a Life Sciences Verification Program developed with the US government, initially limited to US organisations. Anthropic describes it as shipping “the strongest overall cyber capabilities of any model we have released.” The distinction between the two is not capability; it is who Anthropic has decided to trust with it.

Gemini 3.8 Flash is Google’s fast tier at an introductory $0.75 per million input and $3.75 output, rates that double on 1 January 2027 to $1.50 and $7.50. It scores 54.9 percent on HLE-Verified and, by Google’s account, beats most larger frontier models on autonomous software engineering and on legal and finance agent benchmarks. Gemini 3.8 Flash Cyber is the same family tuned for vulnerability work, and Google’s framing is explicitly defender-first: it “prioritised vulnerability fixing from the start” over exploitation. Google’s Chrome security team reports it produces 2.6 times more correct vulnerability patches than the best larger commercial models. It is not in the public API. Access runs through a new Fairwind Program limited to four categories, government authorities, critical-infrastructure operators, software maintainers and core technology platforms, each of which must enforce phishing-resistant multi-factor authentication for every user, restrict the model to internal security, incident-response and penetration-testing teams, and track who used it and how. Reports put the initial partner count above 650. There is no published price because there is no public access.

GPT-6 Astra launched on 3 September at $10 per million input and $50 per million output, the same list price as Fable 5.1. It is the first OpenAI model to meet the “Critical” cybersecurity threshold in the company’s Preparedness Framework. OpenAI’s definition of that threshold is worth quoting in full: a model crosses it if it “can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high-level goal.” During testing, Astra found two previously unknown vulnerabilities in V8, the JavaScript engine inside Chrome, as part of an exploit chain, and scored 100 percent on ExploitBench, which measures whether a model can write working exploits for known vulnerabilities. The production model ships with refusal training and monitoring; the full cyber capability is gated behind a trusted-access programme, with a channel called Daybreak Blue to expand defensive use.

Muse Spark 1.3 from Meta is the opposite kind of release. It is not a capability jump but a price point: roughly $0.10 per million tokens blended, which makes it the cheapest model in the top tier of any leaderboard and sets the floor that every “good enough” workload will be measured against.

Three days, five releases, two access regimesTimeline from 1 to 3 September 2026. 1 September: Claude Fable 5.1, general availability, and Mythos 5.1, trusted access. 2 September: Gemini 3.8 Flash, general availability, and Flash Cyber, Fairwind programme only. 3 September: GPT-6 Astra, general availability with cyber capabilities behind trusted access, and Muse Spark 1.3, general availability at about ten cents per million tokens. Mon 1 SepTue 2 SepWed 3 Sep Claude Fable 5.1$10 / $50 · 1M context · cache −75% Claude Mythos 5.1same weights · verification programmes Gemini 3.8 Flash$0.75 / $3.75 until 31 Dec, then doubles 3.8 Flash CyberFairwind defenders only · no public price GPT-6 Astra · Muse Spark 1.3$10 / $50 · ~$0.10 blended Astra cyber tier“Critical” threshold · trusted access anyone with an API keygated by identity and vetting
Figure 1. Above the line, what you can buy. Below it, what you have to be approved for. Three of the four labs now ship a tier that money alone does not unlock.

The change that matters: capability is now gated by identity

For three years the frontier was a price list. If you could pay, you could call the best model, and the only difference between a bank and a hobbyist was the invoice. This week, three of four labs shipped a tier that cannot be bought. Mythos requires admission to a verification programme. Flash Cyber requires a background check on “security history” and what Google calls “ethical operations,” plus contractual commitments about who inside your organisation may touch it. Astra’s cyber capability requires trusted access. In every case the gate is explicitly about who you are: a government, a critical-infrastructure operator, a security vendor, a lab partner.

This is not a one-off. It is the labs converging on the same answer to the same problem. Each has published a capability framework that names thresholds at which a model becomes dangerous to release openly. Astra crossed OpenAI’s. Anthropic says Mythos 5.1 is its most cyber-capable model ever. Google built Flash Cyber as a separate variant so that the general model could stay general. Once a lab concludes that a model can find zero-days in hardened systems without supervision, it has two choices: hold the capability back from everyone, or release it to the people most likely to use it defensively and least likely to leak it. All three chose the second, and the mechanism they chose is vetting.

Gated tierProgrammeWho is eligibleWhat you must commit toPrice
Claude Mythos 5.1
Anthropic · 1 Sep
Cyber Verification Program; Life Sciences Verification Program (built with the US government)Defensive security teams; life-sciences professionals. US organisations first.Verification of organisation and use; standard safeguards remain for everything outside the programme’s scope.$10 / $50, same as Fable
Gemini 3.8 Flash Cyber
Google · 2 Sep
Fairwind Program, with CodeMenderGovernment authorities, critical-infrastructure operators, software maintainers, core technology platforms; some academic defensive labs. Reported 650+ partners.Background check; phishing-resistant MFA for every user; access limited to security, incident-response and pentest teams; usage tracking per employee.Not published
GPT-6 Astra, cyber tier
OpenAI · 3 Sep
Trusted access programme; Daybreak Blue for defensive expansionInitial group of testers; defenders through Daybreak Blue.Not published in detail. General model ships with refusal training and activity monitoring.$10 / $50 for the general model
Table 1. The three gated tiers shipped this week, from the labs’ own announcements. The pattern is the same in each: capability is unchanged, access is conditional on vetting, and the vetting is about institutional identity and internal controls, not spend.

For an enterprise, this creates a new axis in model strategy that did not exist a month ago. Access is no longer a procurement question. It is a relationship and a compliance question. The groups that will have the strongest models for security work in twelve months are the ones that start the vetting conversations now, with a clear statement of use, a named owner, phishing-resistant authentication already in place for the teams that will use it, and audit controls the labs can inspect. Google’s commitment list is, in effect, the template: if your security function cannot meet it today, that is the project to start.

It also means something uncomfortable about benchmarks. Public leaderboards will increasingly include models that most of their readers cannot run. A comparison that ranks Mythos 5.1 above Fable 5.1 on a cyber benchmark is true and useless to anyone outside the programme. Treat any leaderboard that mixes gated and general tiers as marketing until you hold the keys.

What “Critical” means if you defend things

It is tempting to read OpenAI’s threshold language as a safety disclosure and move on. It is better read as a forecast about your own attack surface. If a model can find two zero-days in V8, one of the most heavily audited codebases in the world, as a by-product of an evaluation, then the cost of finding exploitable flaws in the systems most enterprises actually run, which are not V8, has fallen by a large factor. Astra is gated. The capability is not unique to Astra, and the gate is a delay, not a wall.

Two consequences follow. First, patch cadence becomes a competitive variable. Google’s argument for Flash Cyber is that defenders need a head start, and its evidence is that the model produces verified, deployable patches faster than people do. If your patching pipeline takes weeks from disclosure to deployment, you are now measured against attackers who may take hours. Second, there is an asymmetry that the Mishcon analysis of this summer’s containment failures put precisely: attackers run models with the refusals stripped out, while defenders run models that decline. The gated tiers are the labs’ attempt to correct that asymmetry by handing the less-restricted model to defenders first. Whether that works depends on whether defenders apply.

Pricing is a quarterly variable, not a constant

Astra and Fable 5.1 at $10 and $50 sit at one end of this week. Muse Spark near $0.10 blended sits at the other. One tracker computed the spread across the top fifteen models at 119 times, using Fable 5.1’s blended rate of $11.90 against Muse Spark’s $0.10. On top of list prices sit the mechanics that actually determine a bill: Anthropic’s cache-read cut; Gemini 3.8 Flash’s introductory rate that doubles on 1 January; GPT-5.6 Sol’s promotional $4 and $20 that runs to 21 November; a scheduled Sonnet 5 increase to $3 and $15 that Anthropic cancelled on 11 August and made the $2 and $10 price permanent instead. Three labs now publish rates with expiry dates.

The practical consequence is that cost per task, not list price, is the number to track, and it has to be re-measured every time a lab moves. A workflow whose model name is hard-coded is now carrying budget risk. A workflow whose model is a configuration value behind a router can take a price cut the day it happens.

How to run an evaluation bench that is faster than the release cycle

The only durable advantage a buyer has in a market that ships weekly is a measurement system that is faster than the shipping. Ours is not sophisticated, and that is the point. It has four parts.

  1. A task set drawn from production, not from a benchmark. Thirty to fifty real tasks per workload class, sampled from logs, with the inputs and the accepted outputs. Contract clause extraction, ticket routing, invoice reconciliation, code review on our own repositories. If a model is going to be used for our work, it is measured on our work.
  2. Automatic grading, with a human sample. Exact-match or rubric graders for bounded tasks; a strong model as judge for open-ended ones; and a ten percent human audit of the judge so that the judge is itself measured.
  3. Cost per task recorded next to accuracy. Tokens by type, including cached and batch, multiplied by the rate card in force on the day, attributed to the task. A new model that moves accuracy by a point and cost by half is a bigger event than one that moves accuracy by five points at triple the cost, and the bench has to make that visible.
  4. A cadence measured in hours. The bench runs on release day, unattended, and posts a table by the next morning. If it takes a week, the next release has already landed.

The bench is also the thing that makes the gated-tier conversation concrete. When our security team applies to a verification programme, the application includes the workloads we will run and how we will measure them, because we already do.

What I am doing about this week

  1. Running the bench, not switching. Every model in this week’s batch will have a quieter point release within a month. We evaluate now and migrate when the numbers and the terms are stable.
  2. Moving model names out of code. Where a model identifier still lives in application code rather than in the router’s configuration, that is a defect to fix this quarter.
  3. Starting the verification applications. For a group with regulated subsidiaries, the gated cyber tiers are exactly what our security teams will need. The application processes take months and have prerequisites, phishing-resistant MFA among them, that are worth having regardless.
  4. Budgeting with expiry dates. Every promotional rate goes into the cost model with its end date. The January doubling of Gemini 3.8 Flash is already a line in next year’s plan.
  5. Re-reading the safeguard changes. An 85 percent drop in benign medical refusals changes which workloads are viable in our healthcare businesses. That is a product decision, not a model decision, and it needs a product owner.

Reading the system card as a regulated buyer

Most of the attention on Fable 5.1 went to the benchmarks. For a group with businesses in healthcare, financial services and critical infrastructure, the system card is the document that decides deployment, and it contains three changes that matter more than any score.

The first is the recalibration of refusals. An 85 percent reduction in biology safeguards firing on benign medical queries is the difference between a clinical-documentation assistant that is usable and one that is not. Every medical deployment we have attempted on earlier models has hit the same wall: a model that refuses to discuss dosages, interactions or differential diagnoses with a clinician because the request pattern-matched to misuse. Anthropic has not loosened the safeguard; it has made it more precise, and precision is what regulated deployments need. The same applies to the 60 percent reduction in cyber interventions, which now permit vulnerability identification for defensive work in the general model. A security team no longer needs Mythos to ask a model to explain a CVE.

The second is Enterprise Frontier Safeguards: zero-data-retention running on customer-controlled infrastructure. Every data-protection assessment we have written for a hosted model has had to argue around the provider’s retention window. A configuration in which the provider retains nothing, verifiably, on infrastructure we control, collapses that section of the assessment to a paragraph. It also moves the residual risk from the provider’s logs to our own, which is where we can audit it.

The third is the anti-distillation mechanism, which prevents manual editing of context in multi-turn conversations. This is a safeguard aimed at competitors, not customers, and it will break workflows that depended on rewriting history, including some evaluation harnesses and some agent frameworks that compact context by editing earlier turns. It is an example of a safety feature with an integration cost, and the integration cost lands on the platform team. We found out by running the bench, which is the point of having one.

OpenAI’s equivalent disclosures for Astra are thinner on enterprise configuration and thicker on capability. The two V8 zero-days are the headline, but the operational detail that matters is that Astra ships with monitoring that can stop unauthorised activity. For a buyer, the question is whether that monitoring produces events the customer can see, and OpenAI has not yet said. Google’s Fairwind commitments are the most explicit of the three about what the customer must do, and for that reason the most useful as a template.

A buyer’s calendar for the next quarter

Because the labs now publish prices with expiry dates and ship on each other’s calendars, a model strategy has to be a calendar as well as a decision. Here is ours for the fourth quarter, written this week, with the dates that are already known.

  • By mid-September: bench results for all five models on our twelve workload classes, with cost per task at current rates and at the rates that apply after each promotion ends. Decision on which workloads migrate, and to what.
  • Before 21 November: GPT-5.6 Sol’s promotional rate ends. Any workload still on it is re-costed at list or moved. In practice this is the deadline to decide whether GPT-6 Sol replaces it.
  • Before 31 December: Gemini 3.8 Flash doubles on 1 January. Workloads on it are either locked into a committed rate, if Google offers one, or re-routed. This is the largest single budget event of the quarter for our high-volume tiers.
  • Ongoing: verification applications for the gated cyber tiers submitted by the security function, with the prerequisite controls, phishing-resistant MFA and per-user access tracking, in place before submission.
  • Standing rule: no production migration within thirty days of a release. Point releases, price changes and quiet capability fixes arrive within that window on every model this year.

None of this is strategy in the grand sense. It is calendar hygiene. But the market has made calendar hygiene the thing that separates a platform that benefits from the release cadence from one that is whipsawed by it.

The counterargument, and why it only half holds

The strongest objection to all of this is that gating is temporary theatre. Open-weight models trail the frontier by months, not years. Alibaba’s Qwen3.8-Max, released in August with weights following a week later, scores near or above Opus 5 on many of the labs’ own benchmarks. If the capability that Mythos gates is available in an open model by spring, what has the gate achieved?

Half of that is right. The gates buy time, and the time is finite. But the half that is wrong matters more for an enterprise. The gated tiers are not just capability; they come with the labs’ monitoring, refusal training, incident reporting and, in Google’s case, a patching agent and a programme of partners. An open model with the same raw capability comes with none of that, and the organisation that runs it carries the whole containment problem itself, a problem this summer’s incidents showed even the labs have not solved. The choice is not gated versus open. It is whose controls you are relying on, and whether you have your own.

The pace will not slow. The labs have shown they can ship on each other’s calendars. For a buyer, the only things that compound are a bench that runs in hours, an architecture that treats the model as configuration, and a security function that can pass a vetting process. Everything else is this week’s news.

Sources

  1. Anthropic. Introducing Claude Fable 5.1 and Claude Mythos 5.1. 1 September 2026. System card, 1 September 2026.
  2. Axios. Anthropic releases new models, cuts agent costs. 1 September 2026.
  3. Google DeepMind. Introducing Gemini 3.8 Flash and 3.8 Flash Cyber. 2 September 2026. Fairwind criteria via How AI Works.
  4. OpenAI. Path to Astra: critical capabilities and frontier safeguards. September 2026. The Hill, OpenAI says Astra met critical cybersecurity threshold. Axios, OpenAI to limit access to Astra’s most powerful cyber capabilities.
  5. Local AI Zone. September 2026 AI model updates: a 119× price spread and a $0.10 floor.
  6. Mishcon de Reya. When the sandbox wasn’t a sandbox. 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn