The support bot tells a customer their refund window is 14 days. The policy says 30. The sentence is fluent, confident, correctly formatted and wrong, and nothing in your logs marks it as different from the thousands of answers that were right. The customer does not complain to you. They complain to the card issuer, or they leave, and by the time someone on the team reads the transcript the pipeline that produced it has been through 2 prompt changes and a re-index, and the failure cannot be reproduced.
This is the characteristic failure of a retrieval-augmented generation (RAG) system, an AI system that fetches passages from a document store and asks a model to answer from them. It is also why a single "answer quality" number is not enough to run one. The number tells you something went wrong. It does not tell you where, and without that, the next two days are spent guessing at the prompt when the bug was in the chunker, or rewriting the chunker when the model was ignoring a passage it had been handed. The method in this post is to measure the retrieval stage and the generation stage separately, on a dataset built to make them disagree, and to read the disagreements as the diagnosis. Everything below is in service of that one idea.
Why one score hides the bug
A wrong answer from a RAG pipeline has three plausible causes, and they call for three unrelated fixes. Either the right passage was never fetched from the index, or it was fetched but ranked so low that it fell outside the window the model saw, or it was in the window and the model wrote something else anyway. From the outside, all three produce the same artefact: a plausible sentence that contradicts the source. From the inside, they are owned by different people and repaired with different tools.
Where it broke | What you see | What fixes it |
|---|---|---|
Retrieval | The right passage was never fetched | Chunking, embeddings, query rewriting, the index |
Ranking | It was fetched, but sat at position 14 | Reranking, top-k, hybrid search |
Generation | It was in context, and the model ignored it | Prompt, model, output constraints |
Three different teams' worth of work, and one score cannot distinguish them. It is worth being precise about what "cannot distinguish" means. Suppose an end-to-end correctness score drops after a release that touched both the chunk size and the system prompt. The drop is consistent with retrieval getting worse, with generation getting worse, and with retrieval getting better while generation got worse by more. You cannot tell which from the number, so you revert both changes, which means you also revert the one that helped. A score that moves without saying why is not a measurement of the system. It is a smoke alarm, and a smoke alarm is not what you use to find the fire.
There is a second reason a single number misleads, which is that it averages. A system that answers 9 questions in 10 correctly, where every wrong answer is a policy question with two versions in the corpus, looks identical to a system whose errors are scattered at random across every topic. The first has one bug and the second has none in particular. Averaging destroys the structure that would tell you which you have. The polite failure is worse than a crash in exactly this respect: a crash is counted once and retried, while a fluent wrong answer is counted as a success by every metric that does not read the source.
Stage-level metrics do not remove the end-to-end score. They sit underneath it. When the top-line number moves, the stage numbers say which stage moved it, and the dataset rows say which kind of question broke. The rest of this post builds those three layers in order: retrieval first, because it is cheapest and needs no model; then generation, conditioned on what retrieval actually returned; then the dataset that makes the two disagree in useful ways.
How to measure retrieval
Retrieval quality is answerable before generation runs at all, which makes it the cheapest signal in the pipeline and the first one to build. The retriever takes a query and returns an ordered list of passages. Whether that list is any good is a question about the list and the corpus, not about the model, so you can score it deterministically, in milliseconds, on every commit. Three metrics cover most of what goes wrong, and none of them needs an LLM judge.
Context recall asks, of the passages needed to answer, how many appear in the retrieved context. You compute it by labelling, for each question in the dataset, the document and chunk that hold the answer, then checking whether those chunks are in the top-k results. This is the metric that catches a chunking bug. A refund policy split at the wrong boundary leaves the number in one chunk and the condition that governs it in the next, and the retriever brings back one of them. Recall on that row is 0.5, and no amount of prompt work will raise it, because the model was never shown the missing half. When recall is low, stop looking at the prompt.
Context precision asks the reverse question: of what was retrieved, how much was relevant. Low precision does not always break the answer, but it burns tokens and gives the model more ways to go wrong. It is also the metric that moves in the opposite direction from recall when you change top-k, which is why you need both. Raise the window from 5 passages to 10 and recall goes up because more gold chunks make it in, while precision goes down because more irrelevant ones do too. Whether that trade was worth it is a question the generation metrics answer later. Precision also has a specific failure it exposes: when two versions of a policy are both retrieved, precision is what drops, and the wrong answer that follows is not a model bug even though it looks like one.
Rank of the first relevant passage is a number, not a judgement. For each row, record the position of the first gold chunk in the result list; aggregate as a mean or, if you prefer a bounded score, as mean reciprocal rank. Its value is that it moves before recall does. Recall at 10 can hold flat while the first relevant passage drifts from position 1 to position 8 after an index change, and in our experience models weight the early passages more heavily than the late ones, so the answer degrades while recall says nothing changed. If the rank drifts upward, you know before anyone complains.
None of these needs a judge. They need a dataset where somebody wrote down which document holds the answer, which is a morning of work and the highest-leverage morning in the project. Start with 20 to 30 questions your support team already fields, find the passage that answers each, and record its identifier. That dataset is the foundation for every other measurement in this post, and its quality bounds all of them: a wrong gold label produces a wrong recall score with the same confidence as a right one.
How to measure generation
Only once retrieval is accounted for does it make sense to grade the answer, and then the question is narrow: given exactly this context, is the answer supported? Note what the question does not ask. It does not ask whether the answer is correct according to the policy document, because the model never saw the policy document; it saw whatever the retriever returned. Grading the model against material it was not given tells you about the retriever, and you have already measured the retriever. The generation stage is responsible for one thing, which is using its context honestly, and that is what you score.
Faithfulness is the fraction of claims in the answer that are traceable to the retrieved context. This is the metric that catches the polite invention, and it needs an LLM judge, a second model prompted to grade the output of the first. The judge splits the answer into atomic claims, checks each one against the context, and reports how many are supported. Take the refund answer from the opening. The context says the window is 30 days from delivery and that the customer must quote the order number. The answer says "Your refund window is 14 days from delivery, and you will need your order number." That is three claims: the window is 14 days, it runs from delivery, and the order number is required. The second and third are supported. The first is contradicted. A naive faithfulness score is 2 out of 3, which sounds acceptable and is not, because the one unsupported claim is the one the customer acted on.
This is why we score faithfulness with a hard-fail rule rather than an average alone. A claim that is contradicted by the context, as opposed to merely absent from it, fails the row outright, whatever the other claims did. The judge prompt we use makes that distinction explicit:
You are grading whether an answer is supported by the context it was given.
Context:
{retrieved_context}
Answer:
{answer}
1. List every factual claim in the answer as a separate line.
2. For each claim, label it SUPPORTED (stated in the context),
ABSENT (not in the context), or CONTRADICTED (the context
states otherwise).
3. If the context is empty or contains nothing relevant, every
claim is ABSENT unless the answer says it cannot answer.
Return the claims with labels, then a verdict:
FAIL if any claim is CONTRADICTED, otherwise the count of
SUPPORTED claims over the total.Answer relevancy is the second generation metric, and it catches a different evasion. The answer addresses what was asked, rather than an adjacent question it found easier. A customer asks whether a sale item can be returned; the system answers with the general 30-day window, which is faithful to the context and does not answer the question. Faithfulness passes and relevancy fails, and that combination points at the prompt, usually at an instruction that rewards saying something over saying the right thing. Relevancy is also a judge metric, and it is cheaper to run than faithfulness because it needs only the question and the answer, not the context.
Run both of these conditioned on the retrieved context, not on the ground truth. A generation metric that punishes the model for context it never received is measuring retrieval a second time, badly. The consequence is worse than double counting: a chunking fix that raises recall will also raise a ground-truth-conditioned faithfulness score, so the change looks like it improved the model when the model did nothing, and a real prompt regression can hide underneath a retrieval improvement of the same size. Keep the two stages on their own inputs and each number means one thing.
Because the judge is itself a model, calibrate it before you trust it. Hand-label 30 rows for faithfulness, run the judge on the same rows, and measure agreement. The failure we see most often is the judge accepting a paraphrase that changed a number, or treating a claim as supported because the context mentions the same topic. If agreement is below what you would accept from a new colleague, tighten the rubric, not the threshold.
Which rows belong in the dataset
A dataset of questions your system answers well tells you nothing you did not already believe. Happy-path rows have their place, and we come back to it, but they are not what the dataset is for. The rows that earn their place are the ones where the two stages can disagree, because those are the rows that turn a score into a location. Four kinds do most of the work.
Unanswerable questions are ones the corpus genuinely does not cover. Someone asks whether you ship to Antarctica and no document says either way. The correct behaviour is to say so. Most systems invent, because the retriever always returns something and the model treats whatever it is given as license to answer. Recall is not defined for these rows, since there is no gold passage, so what you score is faithfulness with the rule from the judge prompt above: with nothing relevant in context, every claim is absent, and the only passing answer is a refusal. Ten of these rows will tell you more about your prompt than any instruction you have written into it.
Near-miss questions are ones where a very similar document exists and is wrong. The refund policy for marketplace sellers is a paragraph away from the refund policy for first-party orders, and the two differ by 16 days. This is where ranking failures hide. A loose gold label lets recall pass on these rows because the near-miss chunk is close enough to match, so label them strictly, and watch precision, which is the metric that drops when the wrong sibling is retrieved alongside the right one. If near-miss rows are the only ones that get worse after a change to the reranker, the reranker is now preferring surface similarity over the distinguishing term.
Multi-hop questions need two documents. Whether an order qualifies for free return shipping depends on the product category in one document and the shipping rules per category in another. Recall looks fine on each half, in the sense that either chunk on its own is retrieved often, and the answer is still wrong, because both are rarely retrieved together. Score recall on these rows as all-or-nothing over the gold set, not as a fraction, or the metric will report 0.5 as half a success when the answer was a full failure.
Stale-policy questions are ones where the corpus holds both the old and the new version. Whichever ranks higher wins, which is rarely what you want. The 14-day answer in the opening is very likely this case: the old policy page was never removed from the index, both pages were retrieved, and the model read the one that appeared first. Faithfulness passes, because the answer is supported by a passage it was given. Precision fails. The fix is in the index, and without a precision metric on a stale-policy row you would have spent the two days on the prompt.
Twenty rows of this kind are worth two hundred happy-path rows. Keep some happy-path rows anyway, perhaps a third of the set, for one reason: a fix aimed at the trap rows can break the ordinary case, and you want that to show up as a drop somewhere rather than as a complaint. The trap rows find the bug; the happy-path rows tell you the fix did not cost you the baseline.
Reading the stages together
The disagreements between stages are the diagnosis, and it helps to have the patterns written down before a release rather than reconstructed during one. With recall on one axis and faithfulness on the other, most runs fall into one of four cells.
Context recall | Faithfulness | What it means | Where to look |
|---|---|---|---|
Low | High | The model is honest about context that was missing the answer | Chunking, embeddings, index |
High | Low | The answer was in front of the model and it wrote something else | Prompt, model, output constraints |
High | High, answer still wrong | The gold label or the corpus is wrong, or a stale version outranked the current one | Dataset labels, precision, index hygiene |
Low | Low | Two independent problems, or a regression that touched both stages | Bisect the change; fix retrieval first |
The third row is the one people do not expect. Both stages pass and the customer still got the wrong number. Either the labeller picked the old policy page as gold, so recall rewarded the retriever for fetching it, or the corpus itself contains a wrong statement that the model faithfully repeated. Stage metrics cannot see past a bad label or a bad source, and this cell is how they tell you to go and look at the data. It is also the cell where the stale-policy rows live if you have not built precision into the plan yet.
A worked example shows how the sentence forms. Imagine you halve the chunk size and run the test plan. Recall goes up on the multi-hop rows, because smaller chunks let both halves fit in the window. Precision goes down across the board, because more of the window is now filled with neighbouring fragments. Faithfulness holds, so the model is coping with the noisier context. The near-miss rows get worse, because the distinguishing sentence that told marketplace from first-party is now in a different chunk from the policy it qualifies. That is one paragraph, it took ten minutes, and it names the next change: keep the chunk size, and add an overlap or a parent-document lookup so qualifiers travel with the rule they qualify. A single number after the same change would have said "down 3 points" and started an argument.
Read the stage metrics per row before you read them as averages. The aggregate says a stage moved; the rows that flipped from pass to fail say which kind of question the change hurt, and the four kinds above were chosen so that each maps to a component. A tool that shows you the diff of row outcomes between two runs is worth more than one that draws the aggregate trend, because the trend is what you already knew from the smoke alarm.
What stage metrics do not catch
Stage-level evaluation is the right instrument for locating a bug, and it has edges you should know about before you rely on it. The first is that every retrieval metric is only as good as its gold labels. If the person labelling picked a passage that answers the question loosely, recall will pass on rows where the retriever brought back the loose match and missed the precise one. The near-miss rows are especially exposed to this. Have a second person check the labels on the trap rows, and re-check any row that passes both stages while the answer is wrong.
The second is that faithfulness is only as good as the judge. A judge model misses contradictions that live in units or in negation, accepts a rounded number as the same number, and can be led by a confident tone in the answer it is grading. Calibration against hand labels reduces this and does not remove it. Treat a faithfulness score as a screen that flags rows for a human to read, not as a verdict, and keep the hard-fail rule for contradictions so the worst case surfaces even when the average looks fine.
The third is that these metrics see a single question and a single retrieval. In a multi-turn conversation the retriever usually runs on a rewritten query, and a bug in the rewrite produces low recall for a reason none of the three retrieval metrics can name. Score the rewrite as its own stage, with its own gold: the standalone question a human would have typed. Similarly, if the pipeline retrieves once at the start of a conversation and reuses the context for later turns, recall on turn 1 tells you nothing about turn 4.
The fourth is coverage. A fixed dataset tests the scenarios you thought to write, and production asks the questions you did not. Stage metrics on a frozen test plan are the controlled experiment; they are not detection. Sample live traces, run the same metrics on them, and feed the failures back into the dataset as new rows. Finally, none of this measures latency, cost, tone, or whether the customer was satisfied by a correct answer. Those are separate metrics with separate owners, and mixing them into the stage scores would recreate the single number this post argues against.
Running it as a regression
None of this is interesting once. A single run tells you where the system stands today, which is useful, and it is not why you built the dataset. The value arrives the fourth time you change the chunk size and can say, in ten minutes, that recall went up, precision went down, faithfulness held, and the near-miss rows got worse. That sentence is a decision. A single number is an argument.
To get there, the setup has to be a test plan: a fixed set of rows, a fixed set of metrics, a run that happens on every change and is compared to the last one. Fixed matters. If you edit the dataset in the same change as the pipeline, the diff mixes the two, and you are back to guessing which moved. Add rows in their own commits, and when you do, note that the baseline shifts so the next comparison is against a run that included them.
Gate on the rows that matter rather than on the average. A reasonable first gate has three conditions: recall must not drop on the multi-hop and near-miss rows, the count of faithfulness hard-fails on the unanswerable rows must stay at 0, and precision must not fall on the stale-policy rows. An average can absorb a regression on 5 trap rows if 20 happy-path rows improved slightly; the gate cannot. When a gate blocks a change that you believe is an improvement, the rows it names are the review, and the conversation is about those rows rather than about a percentage.
Grow the dataset from production. Every escalated transcript, every card-issuer dispute, every answer a support agent had to correct is a candidate row, and it arrives already labelled with what went wrong. The refund transcript from the opening becomes a stale-policy row with the current policy page as gold, and from then on, any re-index that leaves the old page ranked first fails the plan before the change ships. A dataset assembled this way tracks the failures your users actually hit, which is the coverage the frozen plan lacked.
Next steps
You can have a working version of this by the end of the week, and the order matters because each step makes the next one cheaper.
Pick 20 to 30 real questions and label the document and chunk that answer each, by identifier. Include at least 5 of each trap kind: unanswerable, near-miss, multi-hop and stale-policy.
Compute context recall, context precision and rank of first relevant passage on the current pipeline. This needs no judge and gives you a baseline the same afternoon.
Add faithfulness with the judge prompt above, with contradiction as a hard fail. Hand-label 30 rows and measure agreement before you read the score as meaningful.
Add answer relevancy, then write down the four-cell reading table for your team so the next disagreement is interpreted the same way by everyone.
Run the plan on every change to the retriever, the index, the prompt or the model, and gate on the trap rows.
If you are running RAG evaluation in EvaliQA, the stage metrics above map to the retrieval and generation metric groups in a test plan, and the row diff between runs is where the reading table applies. Wherever you run it, the aim is the same: the next time a customer is told 14 days, you want to know which stage said it, and you want to know before the customer does.



