On 17 June 2026, Justin Norman, Michael Rivera and Alex Hughes posted the largest evaluation of language-model judges yet published: 21 judges from nine providers, run across MT-Bench, JudgeBench and RewardBench under three protocols, 118 runs, roughly 541,000 individual judgments. The finding is in the title, “Reliability without Validity.” The judges everyone now uses to grade everything agree with themselves almost perfectly and agree with the truth far less than their raw scores suggest. Two production judges reproduce their own verdicts more than 98 percent of the time while preferring whichever answer appears first by a margin wide enough to swing a model selection.
I read it the week it appeared, because we were in the middle of deciding which model, at which effort setting, should run each class of task on the group’s platform, and the instrument we were using to decide was a model judge. This post is about the bench that makes that decision: sampler, task store, runner, graders, human audit, rate card and report. It is a system, not a spreadsheet, and the research tells you fairly precisely where it breaks.
Why the leaderboard does not transfer
A public leaderboard answers a question nobody in a company is asking: which model is best, averaged over tasks someone else chose, graded by a judge someone else validated, at a cost nobody recorded. Three things fail to transfer.
The task distribution. In the Norman protocol MT-Bench contributes 2,391 items, JudgeBench 350 and RewardBench 2,981. None contains a purchase order, a KYC exception, a maintenance log from a steel plant or a customer email in Hinglish. The only distribution whose scores matter is what a company’s agents actually do, and it is not on any leaderboard.
The grader. Norman and colleagues show that on MT-Bench exact-match agreement with human labels overstates discriminative ability by 33.8 to 41.3 percentage points once you correct for chance with Cohen’s kappa. Gemini 3.1 Pro, one of only two judges that stays in the top three on every benchmark, scores 0.849 exact match and 0.511 kappa on the same data. Both numbers are correct. Only one says how much the judge adds over a coin.
The ranking. Across the three benchmarks, judge rankings shift by up to 14 positions, and 11 of the 21 judges move by four or more. Only Gemini 3.1 Pro and Claude Opus 4.6 stay in the top three everywhere. A ranking is a property of the pair, benchmark and judge, not of the models. There is a ceiling effect too: MT-Bench compresses all 21 judges into a kappa spread of 13.5 points, from 0.376 to 0.511, while JudgeBench spreads them across 60.4. When every candidate on your own bench lands within a few points of every other, suspect the bench before the candidates.
The bench as a system
The bench we run has seven parts, each of which exists because something went wrong without it: a log sampler that draws tasks from production traffic; a task store that holds them with labels, class, privacy status and version; a runner that executes each task across a matrix of model, effort setting and prompt version; graders; a human audit that samples the graders; a rate card that prices every run; and a report of three numbers per task class, from which a routing table is edited by hand, never automatically.
The principle is that the bench is a measurement instrument, and instruments are themselves calibrated. The graders are graded by the audit. The rate card is dated. The judge is pinned. The task store has a held-out split that no prompt engineer can read.
Sampling tasks from production
The task store is the bench’s most valuable asset and its largest liability: it is the only thing that makes the bench a measure of our work rather than someone else’s, and it is built from customer data.
Draw from logs, not from imagination. Every task class has a sampler that pulls a fixed number of recent trajectories each week: input, tools called, final output and whatever outcome signal exists, such as a human override or a reopened ticket. Hand-written tasks drift toward what engineers think the system does. Sampled tasks are what it does.
Stratify on what you will route on. If the routing decision is per task class, each class needs a sample large enough to support a decision, not one proportional to traffic. We stratify on class, input length, language and whether the trajectory ended in escalation, and oversample the tail deliberately.
Handle privacy before the store, not after. Items are de-identified before they are written: names and organisations replaced with consistent surrogates, identifiers hashed, free text scanned. Items that cannot be made safe are dropped and the drop rate is reported, because a sampler that silently drops the hardest cases produces a bench easier than production. Access is logged and retention follows the platform’s data policy.
Version the store, and keep a held-out split. Every item carries its sampling date, the model that produced the original trajectory and the label version; a score without a store version is meaningless. A fraction of each class, rotated quarterly, is visible to the runner and the graders and to nobody else. I will explain why below.
Graders, and the arithmetic of agreement
There are three kinds of grader and a fourth thing that is not a grader but keeps the others honest.
Exact match is for tasks with a checkable answer: a label, an extracted field, a tool call with the right arguments, a test suite that passes. It is cheap, deterministic and immune to bias. Use it wherever the task permits, and design tasks so that it permits. A summary cannot be exact-matched; whether the summary mentions the delivery date, yes or no, can.
Rubrics are checklists of binary criteria scored by a model, each phrased so that a careful human could check it in seconds. The item-response-theory work by Junhyuk Choi and colleagues, posted on 31 January 2026 and revised on 29 May, is the best guide I have found to writing them. They treat each judge as a measurement instrument. Their most practical finding is that detailed instructions with explicit rubric definitions cut prompt sensitivity sharply, in one case reducing the coefficient of variation for GPT-4o on dialogue naturalness from 0.28 to 0.10, with chain-of-thought adding a little more. They also found that most judges are less sensitive to quality differences than human raters, that a seven-point scale sometimes made reliability worse than a five-point one, and that “no model showed acceptable consistency across all criteria.” Write rubrics as short lists of yes-or-no checks, not scales, and test each against prompt paraphrases before trusting it.
Model-as-judge, pairwise, position-swapped. For preference tasks that no rubric captures, we use a judge comparing two outputs, and every pair is run twice with the order swapped. Position bias is the absolute deviation of the first-position win rate from one half across the AB and BA runs. Norman and colleagues measured it from 0.002 for Gemini 2.5 Pro to 0.192 for Qwen 3 8B. A judge with position bias of 0.19 that is only ever run in one order is a judge that prefers the left column. Verbosity bias, by contrast, was below 0.011 for all 21 judges under their single pairwise rubric, which differs from the original MT-Bench paper’s warnings about length and from our own experience with absolute-score graders. I will come back to that.
Report kappa, not raw agreement. Every grader is validated against the human audit with Cohen’s kappa, with raw agreement shown beside it so the gap is visible. When Zheng and colleagues reported in 2023 that GPT-4 reached over 80 percent agreement with human preferences, the same level as agreement between humans, the field took it as a licence. Eighty percent on a task where humans also agree 80 percent of the time is a real result. On a task with lopsided labels, where a judge that always picks the majority answer would score 70 percent, 85 percent exact match is modest, and kappa is the number that says so.
The human audit, sized to the decision. A rotating panel of domain reviewers grades a sample of every grader’s output, and the sample size is set by the decision, not by a round number. If the routing question is a two-point gap between two models, the audit needs an interval narrower than that. At 400 audited items and 90 percent agreement, the usual binomial interval is roughly plus or minus three points; at 1,600 it is roughly plus or minus one and a half. Quadrupling the sample halves the interval, and that is the whole budget conversation.
Pin the judge. The judge model, its exact version, prompt, temperature and rubric are recorded with every score, following the Norman protocol of temperature zero and three to five repeated runs per item. Their paper closes with a Minimum Viable Validation Protocol: kappa beside exact match, position bias from AB and BA pairs, at least three test-retest passes, two benchmarks spanning preference and correctness, and a check that position bias stays below 0.10 wherever test-retest exceeds 0.95. A judge change is a bench change and triggers re-validation against the audit set before any routing decision uses it.
Counting the cost
A bench that reports success rate without cost is half an instrument, and the missing half is the half finance reads.
A dated rate card. Every provider’s list prices, per model and token type, including cached input, batch and premium tiers, live in a file with a date and a version. The runner prices every trajectory against the version current on the run date, and the report names it. As of 2 July 2026, the three models in the worked example below are listed at $5 input and $30 output per million tokens for GPT-5.5, $5 and $25 for Claude Opus 4.8, and $1.50 and $9 for Gemini 3.5 Flash, the Gemini output price covering thinking tokens too.
Tokens are not comparable across vendors or versions. Anthropic’s pricing documentation notes that Claude 4.7 and later models use a newer tokeniser that produces approximately 30 percent more tokens for the same text. The per-token price did not change when Opus 4.7 shipped on 16 April; the per-task cost did. A rate card alone would not have caught it. Only cost per task, computed on real trajectories, does.
Effort pinned per class. Each vendor now exposes a knob that trades tokens for quality: a reasoning-effort parameter on GPT-5.5, which defaults to medium; an effort parameter on Opus 4.8 with levels from low to max, defaulting to high; thinking levels from minimal to high on Gemini 3.5 Flash. A task class is routed to a model and an effort, and the bench runs the matrix. One model at high effort against another at low effort is not a comparison.
Tokens-to-done as a first-class metric. The number that matters is not tokens per call but tokens consumed from the start of a task until it reaches an accepted state, including retries, tool calls and whatever the agent does after a grader rejects its first attempt. Tokens-to-done, priced on the rate card, is cost per task. Latency-to-done sits beside it, for reasons that follow.
Cadence and the release-day protocol
The bench runs on two triggers, and neither is “monthly.”
Every model release. Between 16 April and 28 May this year, Anthropic shipped Opus 4.7 and Opus 4.8, OpenAI shipped GPT-5.5 and Google shipped Gemini 3.5 Flash, and each changed something the rate card has to know. GPT-5.5 doubled the per-token price of GPT-5.4, from $2.50 and $15 to $5 and $30. Gemini 3.5 Flash tripled the price of Gemini 3 Flash, from $0.50 and $3 to $1.50 and $9. Opus 4.7 changed the tokeniser. Opus 4.8 cut the fast-mode premium to a third, from $30 and $150 to $10 and $50. A release adopted without the bench is a change nobody measured.
Every prompt change. A prompt edit is a configuration change to the system under test and goes through the same bench as a model change. A pull request touching a prompt bundle carries a bench run identifier in its description, and the merge check fails without one.
The release-day protocol is written down and dull by design. Add the model to the rate card with the date. Pin its effort level per task class, starting from the vendor’s default. Run the held-out split first, so that if the model has seen anything resembling the public split, the gap shows. Run the judge validation before the candidate evaluation, because a new model is often also a candidate judge. Publish the three-number report with the previous champion beside it. Change the routing table only after the audit sample for that class has been graded. Nothing moves to production on release day; what moves is a report.
The three-number report
Each task class gets three numbers, and the discipline is to refuse a fourth. Cost per task, in currency, on the rate card version named in the header, from tokens-to-done. Success rate, from the pinned graders, with each grader’s audit kappa beside it and a confidence interval from the audit sample size. Escalation rate, the fraction of tasks handed to a person, by design or by failure, because that is where the cost of a model’s weakness actually lands. A model with low success and low escalation is confidently wrong. A model with high success and high escalation is expensive in a way the token bill hides.
Who reads it: the platform team reads all three and the audit kappa, weekly. The owner of each task class reads their row and the change since the last run. Finance reads cost per task and escalation rate, monthly, and has learnt to ask for the rate card version. Routing changes are made with the report open and recorded against the run identifier. Nobody reads a leaderboard.
What went wrong building ours
Four failures, each of which the bench now guards against because it once did not.
The judge drifted between versions. Our pairwise judge was referenced by a model alias rather than a pinned version. A provider update moved the alias, and every candidate’s score shifted in the same week. The fix was pinning the judge version, logging it on every score, and a judge regression test: a fixed set of audited pairs the judge must re-score within a tolerance before the bench accepts its output. The Norman finding that two production judges, Qwen 3 8B and Gemini 2.5 Flash, reproduce themselves at 0.992 and 0.988 while carrying position bias of 0.192 and 0.125 is the general form of this. Consistency is not validity.
Evaluation items leaked into prompts. An engineer improving a prompt took the hardest examples from the task store and pasted them in as few-shot demonstrations. The class’s success rate jumped on the bench and did not move in production. The fix was the held-out split, a registry of hashed evaluation items, and a prompt-build check that fails if any bundle contains the hash of a held-out item.
Graders rewarded length. Our absolute-score rubric grader gave longer outputs higher marks on criteria like completeness, and prompts drifted toward longer outputs because the bench preferred them. Norman and colleagues found verbosity bias small under a single pairwise rubric, and our experience agrees; the bias lived in the absolute-score grader. The fix was to move preference grading to position-swapped pairs, add an explicit length-budget criterion to every rubric, and report the correlation between grader score and output length as a diagnostic on every run. When it rises, the rubric is rewritten.
Latency was ignored. The first report had cost and success and no latency column. A model at high effort won a task class on both and was routed to production, where the task turned out to be interactive and the 95th-percentile response time rose past what users would wait for. Users escalated rather than wait, and the escalation rate rose two weeks later. The fix was latency-to-done at the median and 95th percentile, a latency budget per class in the task store, and a rule that a candidate exceeding the budget is disqualified regardless of score.
A worked example
Take one task class: invoice exception triage. The agent receives an invoice, the purchase order it references and the goods receipt, decides whether the three match, and if not, writes a short note classifying the exception for an accounts-payable clerk. The decision is graded by exact match against the clerk’s eventual decision; the note by a four-criterion binary rubric. An escalation is when the agent declines to decide.
The prices below are list prices as of 2 July 2026. Everything else, the token counts, the success and escalation rates and the cost of an escalation, is invented to show the arithmetic. The sample is 400 tasks per candidate at pinned high effort, with 400 audited items, giving an interval of roughly plus or minus two and a half points.
GPT-5.5, reasoning effort high. Say the average trajectory consumes 6,000 input and 1,800 output tokens to done. At $5 and $30 per million, that is 3.0 cents plus 5.4 cents, or 8.4 cents per task. Suppose success is 93.5 percent and escalation 4.0 percent.
Claude Opus 4.8, effort high. Say 7,200 input tokens, higher because of the tokeniser, and 1,500 output. At $5 and $25, that is 3.6 plus 3.75 cents, about 7.4 cents per task. Suppose success is 94.0 percent and escalation 3.5 percent.
Gemini 3.5 Flash, thinking level high. Say 5,800 input and 2,600 output tokens, the output including thinking. At $1.50 and $9, that is about 0.9 plus 2.3 cents, about 3.2 cents per task. Suppose success is 90.5 percent and escalation 6.5 percent.
Read naively, Flash wins on cost by more than two to one, and the other two are a statistical tie on success. Now price the escalation. If a human-handled exception costs $4 in clerk time, again invented, the all-in cost per task is model cost plus escalation rate times $4: about 24 cents for GPT-5.5, 21 cents for Opus 4.8 and 29 cents for Flash. The cheapest model per token is the most expensive per task, and the two models that tied on success are separated by the escalation column. That is the decision the bench exists to make, and it cannot be made from a leaderboard, a price list or a success rate alone. The next run is Flash and Opus at lower effort, because the knob is part of the configuration and the all-in ordering often changes with it.
Rules
- Sample tasks from production logs, stratified, de-identified, versioned, with a held-out split no prompt can see.
- Prefer exact match; where you cannot, use short binary rubrics; where you cannot, use a pinned judge on position-swapped pairs.
- Validate every grader against a human audit with Cohen’s kappa, show raw agreement beside it, and size the audit to the decision.
- Pin the judge’s version, prompt, temperature and rubric, and run a judge regression test before trusting a run.
- Price every run on a dated rate card, measure tokens-to-done, and report cost per task in currency, not tokens.
- Pin effort per task class and evaluate the model-and-effort pair, never the model alone.
- Run the bench on every model release and every prompt change, with a written release-day protocol, and never change routing on release day.
- Report three numbers per class, cost per task, success rate and escalation rate, with latency as a disqualifier.
- When a score moves, suspect the instrument before the model.