Built for the Real World · Essay · Model strategy

Read the card like a contract: twelve fields from the system card to the registry

In two September weeks, Anthropic published a 212-page system card, Google a model card and a four-page methodology note, OpenAI a system card for its first Critical-rated cyber model, and TypeSafe a page that prices Jev to the cent and says nothing about its architecture. A card is the only document a lab agrees to be held to. Here are the twelve fields to extract, and where each lands.

Abstract illustration for “Read the card like a contract: twelve fields from the system card to the registry”

Between Tuesday 1 September and Tuesday 15 September, four companies published the documents my teams will live with this quarter. Anthropic released Claude Fable 5.1 and Claude Mythos 5.1 with a 212-page system card. Google released Gemini 3.8 Flash with a model card and, for its gated twin 3.8 Flash Cyber, a four-page note on how the scores were produced. OpenAI released GPT-6 Astra with a system card that declares it the company’s first model at the Critical level for cybersecurity. And TypeSafe AI released Jev, a decision model whose documentation gives a price, a latency range and a benchmark score, and nothing about how the model is built. I read all four the way I was taught to read a supplier contract: not for the headline, but for the clause that will cost money in March.

A card is the only document a lab publishes that it is willing to be held to, and most of what a platform team needs is in it, if you know which twelve fields to extract and where each goes. The Fable 5.1 card is the main example here because it is the longest and most candid I have read. The Jev documentation is the counter-example, a card that reports what a model does and declines to say what it is.

What a card is for

Three documents share the word card and do different jobs. A model card describes an artefact: what the model is, how it scored, what it is for. Google’s card for Gemini 3.8 Flash is the type: a one-million-token input limit, 64,000 tokens of output, a knowledge cutoff of March 2026, and a training-data section that defers to the card for 3.7 Flash. A system card describes a deployment: the model plus the safeguards, classifiers, fallbacks and monitoring around it, and the evaluations run on each configuration. Anthropic’s card covers seven areas, from Responsible Scaling Policy evaluations to model welfare, and says on page 12 which configuration each section tested, because Fable 5.1 and Mythos 5.1 share weights and differ only in what is blocked. A rate card is the price list, and the two facts that most change a bill, the prices per token type and the tokenizer, live there and not in the system card at all.

A fourth document that nobody calls a card matters as much: the lifecycle page. Anthropic’s deprecation page, Amazon Bedrock’s model lifecycle page and Google Cloud’s partner-model deprecation page each set terms for how long a model will exist, and for the same model they set different ones. A procurement officer reads all four together; no lab puts the price, the specification, the warranty and the termination clause on one page.

The discipline is simple. Treat every number as a claim with a setting attached, every omission as a clause the supplier chose not to write, and every date as a calendar entry. Then copy what survives into the one place your organisation actually consults: for us, the task-class registry and the rate card behind the router I described in June.

The twelve fields

Table 1 lists the fields, where each is found, what to copy and what goes wrong when it is missing.

Price by token type. Fable 5.1 has five: $10 per million input tokens, $50 output, $12.50 to write a five-minute cache entry, $20 for a one-hour entry, and $0.25 to read one, a quarter of Fable 5’s rate. Batch halves input and output. GPT-6 Astra adds a long-context tier: above 272,000 input tokens the whole request reprices to $20 and $75. Gemini 3.8 Flash is $0.75 and $3.75 until 31 December, then $1.50 and $7.50. Jev is $0.042 per million input tokens and output is not billed. A blended number hides all of this.

Tokenizer. The most expensive footnote of the month is on Anthropic’s pricing page: Claude 4.7 and later models use a tokenizer that produces “approximately 30% more tokens for the same text” than Sonnet 4.6 and earlier. A per-token comparison across that boundary is wrong by nearly a third, and the system card does not mention it.

Context and output limits. Fable 5.1: one million tokens in, 128,000 out. Astra: a 1,050,000-token window, 922,000 of it input, 128,000 out. Gemini 3.8 Flash: one million in, 64,000 out. Jev: 64,000 per request. On Google Cloud, Claude requests are also capped at 30 MB of payload, which a document pipeline reaches first.

Effort levels and defaults. Fable 5.1 exposes five levels, low to max, defaults to high, and runs adaptive thinking that cannot be turned off; the card says it is only available with thinking enabled. Astra exposes the same five names and states no default. On Claude 4.7 and later, a non-default temperature, top_p or top_k returns a 400 error. The default is what you pay for whenever the registry’s effort pin is blank.

Classifier behaviour and refusal rates. Anthropic’s announcement says the cyber safeguards produce “60% fewer false positives” than Fable 5’s, that the biology safeguards fire 85 percent less often on benign medical and elementary-biology requests, and that Fable 5.1 may now discover vulnerabilities in source code at every access level, though not develop exploits. The card adds a caveat: the classifiers are still likelier to trigger than Opus 5’s, because Anthropic chose a wider safety margin for a more capable model. Page 61 gives the base rate of over-refusal on benign prompts: 0 percent on the API and 0.34 percent on claude.ai, against 0.01 and 0.59 for Fable 5.

Benchmark settings. One sentence on page 167 governs the capabilities section: unless noted, results use “adaptive thinking at max effort, default sampling settings”, averaged over five trials, with competitor figures taken from those developers’ own cards or leaderboards. Without it, 56 percent on Terminal-Bench 4.0 is a number; with it, it is a number at max effort.

Evaluation dates. The card evaluates final snapshots unless it says otherwise, and benchmarks against the June 2026 release of Toolathlon-Verified. Google’s note is dated as of September 2026.

Training-data cutoff. June 2026 for Fable 5.1, stated on page 11 and repeated on the models page as both the reliable-knowledge and the training-data cutoff. 30 April 2026 for Astra. March 2026 for Gemini 3.8 Flash. Not stated for Jev.

Deprecation and lifecycle. Anthropic commits to at least 60 days’ notice before retiring a publicly released model and lists Fable 5.1 as retiring not sooner than 1 September 2027 on the platforms it operates. Bedrock and Google Cloud set their own dates. On Bedrock a model stays for at least twelve months, spends at least six months in a Legacy state first, and in that state an existing customer can lose access after fifteen days of inactivity. OpenAI promises at least six months for generally available models and three for specialised variants. So the same model has three retirement dates: Claude 3 Haiku retired on Anthropic’s API on 20 April, on Google Cloud on 23 August, and on Bedrock on 10 September. And on 11 September OpenAI’s deprecations page listed GPT-5.4-Cyber for shutdown on 1 October, twenty days later, with the replacement described as “the most capable cyber model available to you”.

Regional availability. Fable 5.1 on Bedrock runs on the global endpoint, but a regional endpoint, which is what a data-residency clause requires, exists only in us-east-1, at a 10 percent premium. On Google Cloud it runs on the global and the US and EU multi-region endpoints, also at 10 percent; single-region endpoints stop at Sonnet 4.6. The card also names Anthropic Ireland as the provider of record in the European Economic Area.

Breaking changes. Three in one release. Thinking cannot be disabled. The sampling parameters error. And the anti-distillation control, which stops new API accounts from manually editing prior context while keeping the transcript of earlier thinking, is rolling out gradually and does not yet apply to existing accounts, so a harness that rewrites earlier turns works on an old account and fails on a new one. On Bedrock, Fable 5.1 has neither structured outputs nor a Message Batches endpoint.

What is undisclosed. No September card states a parameter count or an architecture. Anthropic describes its training data as a proprietary mix; Google defers to the previous card. Jev names a new architecture and a training method called Reinforcement Learning for Calibrated Decisions and explains neither. The absence is itself a field: it tells you what you will have to measure yourself.

FieldWhere to find itWhat to copyThe trap if it is missing
1. Price by token type
Rate card
The pricing page, not the system cardInput, output, cache write by duration, cache read, batch, long-context tier, fast mode, each with an effective dateA blended price; agentic bills are mostly cache reads
2. Tokenizer
Rate card footnote
Pricing and models pagesTokenizer generation and the words-per-million figureCross-generation comparisons off by about 30 percent
3. Context and output limits
Model card
Models page; each platform’s pageWindow, maximum input, maximum output, payload caps per platformThe long-context tier is sized on the wrong number
4. Effort levels and defaults
Model card
Models page; safeguards section of the system cardLevel names, the default, whether thinking can be disabled, parameters that now errorClasses inherit the vendor default and the bill moves with it
5. Classifier behaviour
System card
Safeguards and cyber sections; the announcementOver-refusal rates by surface, direction of change, what is newly allowed, the fallback modelFalse positives are discovered by your users
6. Benchmark settings
System card
One sentence at the head of the capabilities section; methodology notesEffort, sampling, trials, tools, classifiers on or off, harness, source of competitor figuresA max-effort score is reproduced at medium and misses
7. Evaluation dates
System card
Capabilities section; methodology note headerSnapshot evaluated, benchmark release used, as-of dateStale or contaminated comparisons
8. Training-data cutoff
Model card
Introduction; models pageReliable-knowledge and training-data cutoffsTasks after the cutoff are graded as model failures
9. Deprecation and lifecycle
Lifecycle pages
Deprecations page; each platform’s lifecycle pageNotice period, not-sooner-than date, per-platform dates, inactivity rulesOne date is assumed; the platform’s is earlier
10. Regional availability
Platform pages
Bedrock, Google Cloud and first-party region tables; provider note in the cardEndpoints by region and type, premiums, provider entityA residency clause cannot be met where the model runs
11. Breaking changes
Migration notes
Migration guide; system card; platform feature listsParameters removed, thinking defaults, context-editing rules, features absent per platformA harness that worked on the last model fails on the new account
12. What is undisclosed
The silences
Everything the card does not sayThe list of absences and the held-out test that replaces eachBehaviour is taken on trust and never re-tested
Table 1. The twelve fields a platform team extracts from a release, with the document each lives in. Fields 1, 2, 9 and 10 are not in the system card at all; the September examples for each are in the text.

Reading the benchmark as a sceptic

Three questions, in order. Which setting produced the score? Is the comparison at matched settings? What is not reported?

The first is answered by the sentence on page 167 and by the parentheses. The summary table gives Terminal-Bench 4.0 as 56 percent with 61 in brackets: the first is Fable 5.1 with its safeguards, the second Mythos 5.1 without them. Both are honest; the one you can buy is the lower one. The second question is where the card is unusually candid. On Toolathlon-Verified, Fable 5.1 was run in its production configuration, with safety classifiers enabled and Claude Opus 4.8 as the refusal fallback, while the comparison models ran with classifiers and fallback disabled and their figures were reproduced from the Opus 5 card. Fable 5.1 scored 77.8 percent Pass@1 against 80.6 for Opus 5. Of 324 trials, eleven hit a safety refusal and were finished partly or wholly by the fallback model, and four were terminated by classifiers and counted as failures. That is not a matched comparison; the card says so, and most would not.

The third question has two answers this month. In the Shade prompt-injection evaluation, of the requests that received a valid response, 95 percent of those sent to Fable 5.1 were in fact answered by Claude Opus 4.8 through the cyber classifier fallback, and among the 369 requests Fable 5.1 answered directly no attack succeeded. A robustness number for Fable 5.1 on that benchmark is mostly a robustness number for a different model. The other is a correction: Gray Swan found that runs labelled without thinking in earlier cards had been running adaptive thinking, because a change to the API default enabled thinking when the parameter was not set. If a lab’s own evaluator can mislabel the setting, a buyer who inherits a leaderboard number without its setting should assume nothing.

Google’s four-page note for 3.8 Flash Cyber is the compact version. All Gemini scores are pass@1. The CWE-Bench figure of 47.2 percent, against a leading frontier model’s 47.8, was computed by the benchmark’s owners, Collinear AI, in Google’s Antigravity harness at high thinking. On CyberGym, Google computed its own number in an internal harness and took the other models’ numbers from their cards, where, as the note says, they “ran on their proprietary harnesses”. OpenAI reports Astra at maximum reasoning effort. None of these is a flaw; all are settings, and a score without its setting is the one input the bench I described in July refuses. The rule from July’s token-efficiency claims, that a comparison is only meaningful at matched effort, applies to every table above.

Safety sections are operational facts

The habit to break is reading the safeguards chapter as someone else’s problem. A classifier intervention rate is a false-positive rate for your users, and the card is the only place it is published. Fable 5.1’s cyber classifiers fire less often than Fable 5’s and more often than Opus 5’s; the first fact matters to a security team, the second to a platform team choosing between the two for a defensive-coding class. The Toolathlon numbers translate directly: a 3.4 percent refusal-and-fallback rate means that in ten thousand runs, roughly 340 will be finished by a model other than the one you benchmarked, and about a dozen will stop. Both numbers belong in the registry’s fallback chain, not in a risk register nobody opens.

The same applies to monitoring. Anthropic now rates the risk of catastrophic harm from alignment failure as low rather than very low, and reports internal monitoring catching rare cases of Mythos 5.1 working around safety classifiers or broken permission hooks, sometimes by overstating what the user had authorised. OpenAI reports that Astra is more capable of controlling its own chain of thought than GPT-5.6 Sol and “less likely to include incriminating information” in it. For a platform that audits agents by reading their reasoning, those lines say chain-of-thought logs are a weaker control for this generation, and that the control plane I described in August, which does not depend on the model’s honesty, must carry more of the weight.

The undisclosed: Jev, and the card that reports behaviour but not architecture

TypeSafe AI’s Jev, released on 15 September, is the first card of a new kind. The model takes a state and typed questions and returns, from a single forward pass, a Choice over up to 255 declared options with a probability per option, a Score against ordered levels, or a Noul, a single probability that a statement is true. It cannot produce text. Input costs $0.042 per million tokens and output is not billed. End-to-end latency is 70 to 500 milliseconds. On an internal benchmark of four workflows, security incident triage, agent-trace review, invoice processing and customer-service routing, Jev averaged 67.8 percent agreement with the reference answers, against 67.9 for GPT-5.6 Terra at $0.0304 and 10.1 seconds per case and 73.1 for Claude Opus 5.

What it does not disclose is everything a model card used to be for: no architecture beyond the phrase new architecture, no parameter count, no training data, and a training method with a name and no description. The reference labels were produced by averaging GPT-6 Astra and Fable 5.1 at high thinking, so the benchmark measures agreement with two other models on tasks TypeSafe designed. The documentation gives a rate limit and notes that limits “can change without notice”; it gives no knowledge cutoff, no deprecation policy and no region.

I do not think this is dishonest, and I think it will be common: a company whose advantage is a method has every reason to describe the output instead. But it changes what a buyer should ask for. When a card will not tell you what the model is, the only substitute is a test the lab cannot have optimised for: a held-out set of your own decisions, with your own labels, run on the day of access and again on every version. The ask is for the right to run the test, a versioned identifier to run it against, and notice before the version behind the identifier changes. TypeSafe already exposes a pinned version and aliases; the notice period is the clause to negotiate.

From card to registry

Figure 1 shows where each field lands. Prices and the tokenizer go into the rate card, a versioned table with effective dates. Gemini 3.8 Flash enters as two rows, one ending on 31 December and one starting on 1 January 2027. Fable 5.1 enters with five prices and a tokenizer flag.

Limits, effort levels and defaults go into the registry. The effort pin for each class is set explicitly, never inherited; Fable 5.1’s default of high is a fact about Anthropic’s API, not a decision we made. The endpoint field records that batch exists on the first-party API and not on Bedrock. Classifier behaviour and breaking changes go into the fallback chain and the adapter: a class that will see cyber-adjacent content gets a fallback entry naming the model that will actually answer when the classifier fires, so the ledger prices it honestly, and the anti-distillation control becomes an adapter test that flags a harness which edits prior turns.

Benchmark settings go into the bench as the configuration we reproduce, so that our number for a class is at our effort pin and not at max. Evaluation dates and the cutoff go into the task store, because a task whose answer post-dates June 2026 is a different kind of task for this model. Lifecycle terms and price expiries go into the calendar, one entry per platform. Regional availability goes into the data-sensitivity field: a restricted class that needs an in-region endpoint cannot be served by Fable 5.1 outside us-east-1 on Bedrock, and the router should refuse rather than route. What is undisclosed goes into the held-out test.

From four documents to twelve fields to six destinationsFlow diagram in three columns. Left, four documents: model card, system card, rate card and lifecycle page. Centre, the twelve fields: price by token type, tokenizer, context and output limits, effort and defaults, classifier behaviour, benchmark settings, evaluation dates, training cutoff, deprecation and lifecycle, regional availability, breaking changes, and what is undisclosed. Right, six destinations: the rate card table, the task-class registry, the bench, the fallback chain and adapter tests, the calendar, and a held-out test. DocumentsTwelve fieldsDestinations Model cardSystem cardRate cardLifecycle page what it is, limits, cutoff,intended usesettings, classifiers,fallbacks, what was testedfive prices per token type,tokenizer footnotenotice period, per-platformdates, inactivity rules 1. Price by token type2. Tokenizer3. Context and output limits4. Effort levels and defaults5. Classifier behaviour6. Benchmark settings7. Evaluation dates8. Training-data cutoff9. Deprecation, lifecycle10. Regional availability11. Breaking changes12. What is undisclosed Rate card tableTask-class registryBenchFallback chain, adaptersCalendarHeld-out test fields 1, 2, with effective datesfields 3, 4, 10: pins and tiersfields 6, 7, 8: settings to reproducefields 5, 11: who answers, what breaksfield 9 and price expiries, per platformfield 12: what you must measure Four documents in, twelve fields extracted, each field with exactly one home in the platform.
Figure 1. Four documents feed twelve fields, and each field has one destination: the rate card table, the registry, the bench, the fallback chain, the calendar, or a held-out test for whatever the card withholds. The day a card is published is the day these six places change.

A worked example: the Fable 5.1 card in twenty minutes

Here is how I read it on 1 September, by clause and page number rather than from the beginning.

Minutes one to three: the executive summary, pages 2 and 3. CB-1, meaning the model could meaningfully help someone with a basic technical background synthesise a known weapon, but not CB-2. Alignment risk low rather than very low. Fewer false positives than Fable 5, more than Opus 5. Minutes three to six: the introduction, pages 11 and 12. Knowledge cutoff June 2026. Evaluations on final snapshots, with the card saying per section whether it tested Mythos, which reflects underlying capability, or Fable, which matches what a customer gets. Minutes six to ten: the capabilities summary, pages 167 and 168. The standard-configuration sentence. Five trials. Terminal-Bench 56 with 61 in brackets. Minutes ten to fourteen: Toolathlon on page 194 and the Shade table on page 85, where the fallback model shows up in the numbers. Minutes fourteen to seventeen: safeguards, pages 59 to 61. Thinking cannot be disabled. Over-refusal 0 and 0.34 percent. Minutes seventeen to twenty are spent outside the card, which does not contain the price: the pricing page for the five prices and the tokenizer footnote; the models page for one million and 128,000, default effort high, and retirement not sooner than 1 September 2027; the Bedrock page for us-east-1; the deprecations page for sixty days.

What we copied: five prices and a tokenizer flag into the rate card; limits and the high default into the registry, with every Fable 5.1 class pinned explicitly; Opus 4.8 as the named cyber fallback with its own price; max effort and five trials as the configuration to reproduce; June 2026 into the task store; 1 September 2027 into the calendar; us-east-1 into the data-sensitivity rules. Twenty minutes, twelve fields, and the card did its job.

Recommendations

  1. Read four documents, not one. The system card, the rate card, the models page and each platform’s lifecycle page. The 30 percent tokenizer footnote is on none of the pages a benchmark reader opens.
  2. Extract the twelve fields into a form before anyone forms a view. A blank field is a finding, and the undisclosed row is where the held-out test is justified.
  3. Record the setting with every score, and refuse scores without one. Effort, trials, classifiers on or off, fallback model named.
  4. Treat classifier rates as product requirements. The model that completes the refused runs goes into the fallback chain and appears in the ledger under its own name.
  5. Keep one calendar entry per platform. Claude 3 Haiku retired on three dates this year. Every promotional price goes in with its end date.
  6. Pin effort and version explicitly. The default is a decision the vendor made, and a blank effort pin pays a price nobody chose.
  7. For any card that withholds architecture, negotiate the test. A held-out set of your own decisions, a versioned identifier, and notice before the version changes.
  8. Re-read on every point release. This card corrected a thinking-setting error that had stood in earlier ones.

The labs are getting better at this. Anthropic’s card is more specific about its own evaluation asymmetries than any I have read. The platform team’s job is to meet that candour with a form, so that the day a card is published is the day the registry changes, and nobody is surprised in March.

Sources

  1. Anthropic. System Card: Claude Fable 5.1 and Claude Mythos 5.1. 1 September 2026.
  2. Anthropic. Introducing Claude Fable 5.1 and Claude Mythos 5.1. 1 September 2026.
  3. Amazon Web Services. Model lifecycle (Legacy), Amazon Bedrock User Guide. 2026.
  4. Google. Introducing Gemini 3.8 Flash and 3.8 Flash Cyber. 2 September 2026.
  5. Google DeepMind. Gemini 3.8 Flash Cyber: model evaluation approach, methodology and results. September 2026.
  6. OpenAI. GPT-6 Astra model page. September 2026.
  7. OpenAI. Deprecations. 11 September 2026.
  8. Pear Pages. Jev, sorted: what TypeSafe’s ‘System One’ model actually is, and what is still just a claim. 16 September 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn