Built for the Real World · Essay · Retrieval

Grounding is the product: what GaRAGe shows about enterprise RAG

A year after Amazon’s GaRAGe benchmark showed the best model grounding only 60 percent of its answers in relevant passages, and deflecting on 31 percent of the questions it should have, three 2026 papers have found the same shape. For enterprise retrieval the lesson is that the model is the component you swap. The grounding is the product.

Abstract illustration for “Grounding is the product: what GaRAGe shows about enterprise RAG”

A year ago this month, five researchers at Amazon AGI presented a benchmark at the Findings of ACL 2025, held from 27 July to 1 August, that asked a narrower question than most evaluations of retrieval-augmented generation. Not whether a model’s answer was correct, but whether it was built from the passages that were actually relevant, whether it cited them, and whether it declined to answer when the evidence was not there. The benchmark is GaRAGe, first posted to arXiv on 9 June 2025 by Ionut-Teodor Sorodoc and colleagues. It has 2,366 questions and more than 35,000 passages, 4,752 from private document collections and 30,599 from the Web, all labelled by professional annotators. The best model on the strict measure, Amazon’s own Nova Pro, scored 60.67 percent; Gemini 1.5 Flash scored 59.43. No model attributed its claims to the right sources with an F1 above 58.9 percent. When the right answer was to say there was not enough evidence, the best model, GPT-4o, said so 31.1 percent of the time.

I have a personal reason to keep returning to this paper. In section 2.1, step 4, the authors describe how they filtered the questions a model had generated before any human saw them, and the first tool they name is “a classifier to identify well-formed queries”. The footnote points to a model I published on Hugging Face in March 2022 under the name query_wellformedness_score. The way they used it is the first lesson of this post. The larger argument is that in a retrieval system the model is not the product. The grounding is. Three papers from this spring and one from yesterday say the same thing, and most enterprise systems are still evaluated as if it were not true.

A small model in a footnote

In 2018 Manaal Faruqui and Dipanjan Das at Google published a dataset of 25,100 questions from the Paralex corpus of WikiAnswers paraphrases, each rated by five crowdworkers as well formed or not, the ratings averaged into a score between zero and one. The paper, Identifying Well-formed Natural Language Questions, appeared at EMNLP 2018 and on arXiv on 28 August 2018. The task: given a query, decide whether it is a grammatical, complete, explicit question or a bag of keywords with typos. Their classifier reached 70.7 percent test accuracy.

I fine-tuned a RoBERTa classifier on it and published it in March 2022 under an Apache 2.0 licence, because I needed a cheap front gate for user queries and assumed a few other people would too. The model does one thing: it takes a sentence and returns a score, higher for a well-formed query. It penalises broken grammar, missing words and odd casing, and knows nothing about the topic. I thought of it as a utility; the download counts say other people treat it as infrastructure: it peaked at about 8.5 million downloads in a month and has passed 65 million in total. I do not know most of the places it runs. A classifier trained on 25,100 forum questions became one of the most-used query models on the hub because the problem it solves sits at the front of every retrieval system.

That front position is the point. Before a retriever sees a query, something has to decide whether the query is a question at all. “capex approval pune plant april change?” retrieves something; so does “What is the capital expenditure approval threshold for the Pune plant after the April update?” They do not retrieve the same thing, and a system that treats them alike will be judged on the second and fail on the first. A well-formedness score is the cheapest signal in the pipeline: milliseconds on a CPU, a number between zero and one, a decision rule. Below a threshold, rewrite or ask for more; above it, proceed. Every retrieval system has an implicit version of this gate. Very few have an explicit one with a number attached.

GaRAGe used it in that spirit, on the benchmark’s own questions. The authors generated candidates with a model in four steps: plan Web searches from seed topics, run them, generate questions that blend information from several retrieved snippets, then filter. The filter has three parts: the well-formedness classifier, a named-entity check with spaCy that drops questions containing no entity, and deduplication with a sentence-embedding model. Only questions that passed went to human annotators, who validated them again. Before Amazon trusted a question enough to pay professional annotators to answer it, it checked that the question was well formed. The benchmark applied to its own inputs a gate that most production systems do not apply to a user’s.

What GaRAGe measures that your evaluation does not

Most enterprise RAG evaluations measure two things: did the answer match a reference, and did any retrieved passage support it. GaRAGe’s contribution is to make the second question precise. For each of the 35,351 passages, an annotator first decided whether it was relevant at all and, if so, gave it one of four labels: answers the question, related information, outdated, or unknown value. Across the benchmark, 31.2 percent of passages answer the question, 26.4 percent are related, 7.8 percent are outdated, 8.0 percent are unknown, and 26.6 percent are irrelevant. The human reference answer was written only from passages labelled relevant, with a citation marker on each claim, and averages 192 words with about five cited sources.

Three metrics fall out of that annotation.

Relevance-Aware Factuality. The usual factuality score asks whether every sentence in an answer is supported by some passage in the context. GaRAGe asks whether it is supported by a relevant passage, and combines that with an eligibility judgement against the human answer. RAF is the share of answers that pass both tests, and the gap between factuality and RAF measures how much of the answer came from passages a human had marked irrelevant, outdated or unknown. For Gemini 1.5 Flash the factuality score is 70.50 and the RAF is 59.43; for GPT-4o, 59.30 and 52.88; for Claude 3.5 Sonnet, 64.67 and 48.91. Every model loses points, which the authors read as models summarising whatever is in the window rather than selecting what is relevant now.

Attribution. Because the human answers carry citations, the paper scores a model’s citations as precision and recall against them. F1 lands between 48 and 59 for every model, with Claude 3.5 Haiku and Nova Micro top on 58.9. Precision and recall are more informative: the small models cite generously (Nova Micro recall 75.8, precision 48.2), the larger ones cautiously (Nova Pro recall 49.6, precision 56.9). Neither is what a compliance officer wants: a citation on every claim and none to a passage that does not support it.

Deflection. For 427 questions the retrieved passages were judged insufficient and the reference answer is a refusal. The deflection true-positive rate is how often a model refused on those questions. GPT-4o managed 31.1 percent, Gemini 1.5 Flash 27.2, Claude 3.5 Sonnet 25.3, Nova Pro 18.0, with Nova Micro, Nova Lite, Mistral and Mixtral in single digits. False positives, refusing when the evidence was there, stayed below 3.5 percent for every model. The asymmetry is the finding: given weak evidence, a model will almost always answer anyway.

The judge is GPT-4o at a temperature of 0.2, with the prompts published, and the authors checked it against a human pass on 300 answers in which 97 percent of 2,340 claims were grounded in at least one passage. Figure 1 lays the three tables side by side.

Relevance-aware factuality

Percent of answers, GaRAGe Table 2, ACL 2025

Nova Pro
60.67
Gemini 1.5 Flash
59.43
Qwen2.5 32B
52.90
GPT-4o
52.88
Qwen2.5 14B
52.70
Claude 3.5 Sonnet
48.91
Nova Lite
45.97
Nova Micro
37.16
Claude 3.5 Haiku
36.90
Mistral
34.32
Mixtral
34.12

Attribution F1

Against human citations, percent, Table 4

Claude 3.5 Haiku
58.9
Nova Micro
58.9
Claude 3.5 Sonnet
58.6
GPT-4o
58.4
Gemini 1.5 Flash
55.5
Nova Pro
53.0
Mistral
51.0
Mixtral
51.0
Qwen2.5 32B
50.8
Qwen2.5 14B
48.7
Nova Lite
48.1

Deflection true positives

On 427 questions with insufficient grounding, percent, Table 3

GPT-4o
31.1
Gemini 1.5 Flash
27.2
Claude 3.5 Sonnet
25.3
Qwen2.5 32B
21.5
Nova Pro
18.0
Qwen2.5 14B
17.1
Claude 3.5 Haiku
15.2
Mixtral
9.4
Nova Lite
7.5
Nova Micro
7.0
Mistral
5.2
Figure 1. Eleven models on GaRAGe’s three strict measures, from Tables 2, 3 and 4 of the paper (Findings of ACL 2025, GPT-4o as judge); best in each panel in clay. No model clears 61 percent on relevance-aware factuality, 59 on attribution F1 or 32 on deflection, and the leader differs on each.

Why enterprise RAG fails in production

The main table is the headline, but the analysis section is where the enterprise lessons are, because the paper slices its results along exactly the dimensions that break production systems.

Sparse private corpora. Of the benchmark’s passages, 13.4 percent came from four collections standing in for private knowledge bases: Enron email, arXiv abstracts, AWS DevOps guides and SEC filings. Questions over the Enron collection averaged 47.8 percent relevant passages in their retrieved set; general Web questions averaged 85.6 percent. That difference cost more than ten RAF points for every strong model: Nova Pro scored 64.4 on Web questions and 52.3 on private ones; GPT-4o 57.8 and 41.8; Claude 3.5 Sonnet 53.2 and 39.1. The models did not get worse. The context got thinner, and they kept writing at the same length and confidence.

Time-sensitive questions. About a sixth of the questions, 17.5 percent, were labelled fast-changing, meaning the answer depends on information from the previous seven days, and each passage carries its age relative to the question where a date was available. RAF on those questions ran about ten points below the slow-changing and static ones: Nova Pro 51.7, Gemini 1.5 Flash 49.8, Claude 3.5 Sonnet 46.7, GPT-4o 45.7. Deflections cluster there too, because that is where retrieval most often returns something that was true once. An enterprise corpus is a fast-changing corpus behind a slow-changing index: policies are revised, prices change, and the chunk that answered last quarter still ranks first.

Over-summarising. The paper divides questions by the share of relevant passages in their context: high, above two thirds; medium; and low, below a third. RAF drops about ten points from high to medium and a further twenty from medium to low. Tail-topic questions, 55.1 percent of the benchmark, average 44.2 RAF against 50.3 for head topics. The mechanism is the same in each slice: given a mixture of good and bad passages, the model composes a fluent paragraph from the mixture. Nothing in its training rewards it for leaving a passage out.

I recognise all three from our own deployments. The failures are rarely in the model. They are in a corpus thinner than anyone admitted, an index with no notion of time, and a generator that treats its context as a reading list rather than as evidence of varying quality.

A reference pipeline, with a number at every stage

The response to a benchmark like this is not to pick the model at the top of the table; the ordering changes with the slice, and the top score is 60 percent. It is to build the pipeline so that each failure mode is caught at the stage that causes it, and to measure each stage separately. Figure 2 is the shape we use.

Query gate. Score the incoming question for well-formedness and log the score. Route anything below a threshold to a rewrite before retrieval, and track the share of traffic that falls below it. That share is a product metric: it tells you how your users actually type.

Rewrite and decompose. GaRAGe’s own retrieval decomposed each question into sub-queries with a model, then dropped sub-queries whose embedding drifted too far from the original. Do the same, and measure the drift. A rewrite that changes the intent surfaces later as a confident answer to a question nobody asked.

Retrieval with freshness. Store a date on every chunk, and pass it to the ranker and the generator. Measure the relevant share of the top K against a labelled set, by corpus, and the age distribution of what is retrieved for time-sensitive questions. The paper’s 47.8 against 85.6 split is the number that, per corpus, tells you where to spend on ingestion rather than prompts.

Grounding check. Before generation, classify each retrieved passage as answering, related, outdated or irrelevant, with a small model or a cheap call, and drop or mark the last two. Measure RAF on your evaluation set, and the gap between plain factuality and RAF, which is the share of answers written from passages you would not want quoted.

Attribution. Require a citation marker on every claim-bearing sentence, and score precision and recall against reference citations. Watch the two numbers separately, because the paper shows models trade one for the other.

Deflection. If the grounding check finds no passage that answers the question, do not generate. Return a deflection that says what was searched and not found, and route to a person or another corpus. Measure true positives on a deflection subset and false positives on the rest. A deflection rate of zero is a defect, not an achievement.

Six-stage retrieval pipeline with a measurement at each stageSix boxes in a row, query gate, rewrite, dated retrieval, grounding check, attribution, and answer or deflect, each with the metric to track beneath it, a dashed path from the grounding check to deflection when no passage answers, and a note on where stale, sparse and over-summarised failures surface. Six stages, six numbers: the pipeline we use for enterprise retrieval 1 Query gate 2 Rewrite 3 Retrieve, dated 4 Grounding check 5 Attribute 6 Answer or deflect no passage answers: skip generation, deflect Measure well-formedness score share below threshold Measure sub-query drift from the original intent Measure top-K relevant share passage age at query Measure RAF, and the gap to plain factuality Measure citation precision and recall vs human Measure deflection true and false positive rates Where the three failure modes surface Stale corpus: stage 3 passage age, stage 4 outdated labels Sparse corpus: stage 3 relevant share, stage 6 deflections Over-summarising: stage 4 RAF gap, stage 5 citation precision
Figure 2. The reference pipeline. Each stage emits one number, and the three failure modes GaRAGe isolates, stale, sparse and over-summarised, surface at specific stages rather than in the model choice.

Relevant is not warranted: the 2026 evidence

The year since GaRAGe has sharpened the point. On 7 May 2026, Hailey Onweller and colleagues posted Cited but Not Verified, which checks the citations in deep-research agents’ reports for whether the link resolves, the page is on topic, and the page supports the claim. Frontier models kept link validity above 94 percent and relevance above 80 percent, supported their claims between 39 and 77 percent of the time, and lost about 42 percent of fact-check accuracy on average as retrieval calls rose from two to 150. More searching produced more citations and less support.

On 27 May, Pin Qian and colleagues posted Relevant Is Not Warranted, which names the failure “citation laundering”: a real, relevant source presented as warrant for a stronger claim than it supports. Their FORCEBENCH holds a cited passage fixed and pairs a claim calibrated to the evidence with a version pushed harder along one of five axes: relation, modality, scope, temporal validity or numeric specificity. A good evaluator should prefer the calibrated claim. On a 198-pair set, token and entity overlap metrics got the order wrong on 32.8 to 36.4 percent of pairs, and four model judges with a generic support prompt had an aggregate violation rate of 47.2 percent. Prompting explicitly for warrant strength brought that to 24.5 percent, better and still wrong a quarter of the time.

And yesterday, 15 July, a working-notes paper from the DS@GT ARC team for the CLEF 2026 LongEval task reported that frontier models identified the relevant documents and then did not use them when composing the answer, scoring well on fluency and poorly on grounding, while a pipeline that filtered chunks before generation and enforced entailment of each claim to its citation afterwards improved grounding modestly. Three groups, three settings, one shape: a relevant citation is not a supported claim, and the metrics most teams use cannot tell the difference.

One question, three ways to fail

Here is a composite from our own work, details changed. A finance analyst in a subsidiary asks the assistant: “What is the capital expenditure approval threshold that the Pune plant can sign off locally, and did it change after the April delegation-of-authority update?”

The query gate passes it. The rewrite produces two sub-queries, the current threshold and the April change, both close to the original. Retrieval returns eight passages: the April 2026 circular, which contains the new threshold; the 2023 version of the same policy, which contains the old one; an internal FAQ that explains what delegation of authority means; an email thread in which two managers discuss an exception; a board slide about capital planning; and three passages about an unrelated plant.

The first failure is staleness. Without dates, the 2023 policy ranks above the April circular because it is longer and uses the question’s words more often, and the model quotes the old threshold with a citation that is real, relevant and wrong. A date on every chunk and a freshness term in the ranker catch this; so does a grounding check that labels the 2023 passage outdated. This is the paper’s fast-changing slice in one question.

The second failure is sparseness. If the April circular was never ingested, because it lives in a mailbox and not in the document system, the retrieved set has no passage that answers the question. A system without a deflection path answers anyway, from the FAQ, the email and the slide, with a confident number that nobody wrote. The right output is a deflection that names what was searched and points to the policy owner. This is the paper’s private-corpus slice, where relevant context is around half the Web level and the models keep writing.

The third failure is over-summarising. Even with the circular present and ranked first, a generator given all eight passages will fold the FAQ’s definition, the email’s exception and the slide’s projection into one smooth paragraph, each sentence with its own citation. The citations will pass a link check and a relevance check, and the paragraph will say more than the circular warrants: the FORCEBENCH failure. A grounding check that passes only passages labelled as answering, and an attribution check that scores precision, catch it.

Three failures, three stages, three numbers. None of them is a model choice.

Building the evaluation set from your own documents

The benchmark will not transfer, but the method will. Sample questions from your own logs rather than writing them, and gate them as GaRAGe gated its generated ones: well-formedness score above a threshold, at least one named entity, deduplicated by embedding. Run your production retriever and keep the top K passages exactly as the generator would see them, with source and date. Have annotators label every passage relevant or not, and the relevant ones as answering, related, outdated or unknown; have them write a reference answer only from the relevant passages, with a citation marker on each claim, and a deflection where nothing answers. Tag each question for time-sensitivity, complexity and corpus. GaRAGe paid its vendor 9.65 dollars per sample, then had an independent pass check 300 answers. A few hundred questions done this way will tell you more than ten thousand scored against a reference answer alone.

Then score with an LLM judge, starting from the paper’s published prompts, and calibrate it against a human-labelled slice before you trust it. Report RAF, attribution precision and recall, and deflection true and false positives, each by corpus and by time-sensitivity. The slices tell you which corpus to re-ingest, which ranker needs a date, and which model, if any, to switch.

Recommendations

  1. Put an explicit, scored query gate at the front of every retrieval system, and track the share of traffic below the threshold as a product metric.
  2. Date every chunk at ingestion, and give the ranker and the generator the date. Treat an undated corpus as a defect; the ten-point gap on fast-changing questions is a freshness problem first.
  3. Classify retrieved passages before generation, and keep outdated and irrelevant ones out of the window or marked. Measure the gap between factuality and relevance-aware factuality.
  4. Build a deflection path and a deflection subset. A system that never says it does not know is hallucinating on a schedule you cannot see.
  5. Score attribution as precision and recall against human citations, not as the presence of a link.
  6. Build your own GaRAGe: a few hundred real questions, every passage labelled, reference answers written only from relevant passages, sliced by corpus and by time-sensitivity.
  7. Spend on ingestion before prompts. The paper’s private-corpus gap is a coverage problem that no prompt closes.
  8. Choose models last, on your own slices, after the pipeline is instrumented. The top of GaRAGe’s table is 60 percent, and the leader changes with the slice.

The model of mine in this paper’s footnote is a reminder of how little of the work is the model. It is a small classifier and a threshold, and it does its job because it sits in the right place. Grounding is the product. The model is the component you can swap.

Sources

  1. ACL Anthology. GaRAGe: A Benchmark with Grounding Annotations for RAG Evaluation, Sorodoc, Ribeiro, Blloshmi, Davis and de Gispert, Findings of ACL 2025, 27 July to 1 August 2025. arXiv, 2506.07671, 9 June 2025.
  2. GitHub. amazon-science/GaRAGe, dataset and annotation schema, CC BY-NC 4.0.
  3. Hugging Face. Ashishkr/query_wellformedness_score, model card, published 2 March 2022.
  4. Hugging Face. google-research-datasets/google_wellformed_query, dataset card.
  5. arXiv. Faruqui and Das. Identifying Well-formed Natural Language Questions, EMNLP 2018, 28 August 2018.
  6. arXiv. Qian et al. Relevant Is Not Warranted: Evidence-Force Calibration for Cited RAG. 27 May 2026.
  7. arXiv. Onweller et al. Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents. 7 May 2026.
  8. arXiv. Michaels and Johnson. DS@GT ARC at LongEval: Citation Integrity and Factual Grounding in Scientific QA, CLEF 2026 LongEval working notes. 15 July 2026.
Ashish Kumar

Ashish KumarHead of AI & Data Platform at Tata Group. Previously applied AI at Ola Krutrim, data science at Salesken, and conversational AI at Reliance Jio Haptik and Active.Ai. Full biography · LinkedIn