A retrieval-augmented generation (RAG) system gives a wrong answer and you have a score that says so. The score is not the problem. The problem is that the answer was produced by two components in sequence, a retriever that fetched passages from your knowledge base and a generator that wrote an answer from them, and the score does not say which one failed. If the retriever brought back the wrong passages, no amount of prompt engineering will fix the answer. If the retriever brought back the right passages and the model ignored them, re-indexing the knowledge base will change nothing. Teams that only track an end-to-end correctness number spend weeks fixing the wrong stage because the number cannot tell them which stage to look at.
This post is about the metrics that fix that. Some RAG metrics read only the retrieved passages and score the retriever. Some read the final answer against the passages and score the generator. On their own each one answers a narrow question. Read in pairs, they localize the bug: one combination means retrieval missed the facts, another means retrieval delivered and the model drifted, a third means the model answered correctly from its own memory while retrieval was returning noise. We will go through the retrieval metrics, the generation metrics, the pairs that matter, and the setup mistakes that make all of them lie.
Why one score cannot localize a RAG bug
An end-to-end metric compares the final answer with something: a gold answer, the question, a rubric. In EvaliQA that is Answer Precision (the answer against an expected output) or Answer Relevancy (the answer against the question). These are the right metrics for the question "is the product good enough to ship", because they measure what the user sees. They are the wrong metrics for the question "what do I fix", because the answer the user sees is the product of two decisions made by two different components with different failure modes and different owners.
Consider the two ways a RAG answer goes wrong. In the first, the retriever returns passages that do not contain the fact the user needs: the index is stale, the chunking split a table across two chunks, the embedding model does not understand the domain vocabulary, or the top-k cutoff dropped the one relevant chunk. The generator then either admits it cannot answer, which fails correctness, or fills the gap from its own training data, which may or may not be right. In the second, the retriever returns exactly the right passages and the generator misreads them, over-summarises, picks the wrong number from a table, or adds a claim the passages do not support. Both paths end in a failed Answer Precision row. The row looks identical in the run's pass rate.
The fix for the first path is retrieval work: re-chunk, re-embed, tune the ranker, raise k. The fix for the second is generation work: tighten the prompt, swap the model, lower the temperature, change how context is presented. These are different teams in most organisations and different weeks of effort in all of them. A metric set that cannot assign the failure to one of the two paths is, for debugging purposes, a metric set that cannot tell you anything beyond "look harder".
The way out is to measure each stage against its own input. The retriever's job is to turn a question into passages that could answer it, so retrieval metrics read the question, the passages, and optionally a gold answer, and never look at what the model wrote. The generator's job is to turn passages into an answer that stays inside them, so generation metrics read the answer against the passages. When each stage has its own score, the combination of scores on a failing row points at the stage that broke. The rest of this post is about those scores and those combinations.
The four columns the metrics read
Every RAG metric in EvaliQA reads some subset of four columns on a dataset row, and knowing which metric reads which column is most of what you need to interpret a score. The input column is the user's question. The retrieval_context column holds the passages the retriever returned for that question, in the order it returned them. The actual_output column is the answer the AI system produced, filled in at run time. The expected_output column is the gold answer you wrote in advance, and it is optional: some metrics need it and some do not. The Datasets methodology page frames this as three things every row must carry: what goes in, what good looks like, and anything the system needs to do the job, and for a RAG plan the passages are that third thing.

Which columns a metric reads tells you what it can and cannot see. A metric that reads input and retrieval_context only is scoring the retriever and is blind to the answer; it will give the same score whether the generator wrote a perfect answer or gibberish. A metric that reads actual_output and retrieval_context is scoring whether the generator stayed inside the passages, and it is blind to whether those passages were the right ones. A metric that reads actual_output and expected_output is scoring the end result and is blind to both stages. None of this is a limitation to work around; it is the mechanism that makes localization possible, because each blind spot is exactly what another metric covers.
The two columns that need the most care are retrieval_context and expected_output. The retrieved passages have to be the passages the model actually saw on that run, not a snapshot from the day you built the dataset, otherwise a generation metric is comparing the answer with context the model never received. EvaliQA handles this on the connector: the connector's output settings include a field called Retrieval context path, which points at where the retrieved passages live in your AI system's response, so the column is filled at run time from what the system really retrieved. The gold answer has to be written so that its claims can be checked against passages, which means short, factual and specific rather than a paragraph of prose. Both of these come up again when we get to setup mistakes.
Retrieval metrics
EvaliQA ships three metrics that score the retriever, all in the RAG category of the test plan wizard and all documented on the RAG and general LLM-judge metrics page. They share a column, retrieval_context, and differ in what they compare it with. Each uses the plan's judge model, which is set once on Step 3 of the wizard and applies to every LLM-based metric on the plan.
Contextual Relevancy reads input and retrieval_context and asks whether the retrieved passages are on topic for the question. It judges the passages themselves: did retrieval bring back material that could plausibly answer the question, or noise. It needs no gold answer, which makes it the retrieval metric you can run on any row, including rows promoted straight from production traces. Its default threshold is 0.6. What it does not catch is a retriever that returns relevant-sounding passages that happen to miss the specific fact; a chunk about Faithfulness thresholds is on topic for a question about Faithfulness thresholds even if the number the user needs was in the next chunk.
Contextual Precision reads actual_output and retrieval_context and scores precision at k weighted by relevance, which in plain terms is whether the retriever ranked the most useful chunks first. A score of 1.0 means the top chunks are the most relevant ones; 0.0 means the order is inverted. The default threshold is 0.7. Order matters more than it sounds: if you retrieve 20 chunks and truncate to the top 5 before the prompt, or if the model pays more attention to the start of its context than the end, a relevant chunk ranked eighth is as good as absent. The metric has an optional top_k parameter that caps the number of chunks evaluated, and you should set it to the number the model actually sees so the score reflects reality rather than the full retrieval list. Contextual Precision is the metric for teams working on the retriever itself. If retrieval is a black box you do not own, or you feed every chunk to the model without ranking, the docs are explicit that it will only produce noise.
Contextual Recall reads input, expected_output and retrieval_context and scores the proportion of claims in the gold answer that the retrieved passages support. A score of 1.0 means retrieval brought back everything needed to produce the expected answer; anything less means at least one needed fact was missing from the context. The default threshold is 0.7. This is the most diagnostic of the three because it measures the thing the generator depends on: not "were the passages related" but "could a perfect generator have written the right answer from these passages". Its cost is that it needs a gold answer on every row, so it cannot run on unlabeled production rows.
Metric | Reads | Question it answers | Needs a gold answer | Default threshold |
|---|---|---|---|---|
Contextual Relevancy | input, retrieval_context | Are the passages on topic for the question | No | 0.6 |
Contextual Precision | actual_output, retrieval_context | Are the most useful chunks ranked first | No | 0.7 |
Contextual Recall | input, expected_output, retrieval_context | Did retrieval bring back every fact the gold answer needs | Yes | 0.7 |
The three overlap enough that you rarely want all of them on one plan. The docs recommend skipping Contextual Relevancy if you are already running Contextual Precision and Contextual Recall. A reasonable split is Contextual Recall when you have gold answers and want to know whether retrieval delivered the facts, Contextual Precision when you are tuning the ranker, and Contextual Relevancy when you have no gold answers and need some retrieval signal anyway.
Generation metrics
The generator's job is to write an answer that is correct and that stays inside the retrieved passages. Those are two different properties and EvaliQA scores them with different metrics. Correctness is Answer Precision; grounding is Faithfulness; and a third metric, Answer Relevancy, catches the case where the answer is correct and grounded but not actually about the question.
Faithfulness reads input, actual_output and retrieval_context and checks whether every claim in the answer is supported by the passages. A score of 1.0 means all claims are grounded; anything less means at least one claim was fabricated or brought in from outside the provided context. This is the RAG-specific hallucination metric, and the word hallucination here has a precise meaning: not "the answer is wrong" but "the answer says something the passages do not say". An answer can be factually correct and still fail Faithfulness if the fact came from the model's training data rather than the context, and that distinction is the whole point. The default threshold is 0.7, and the docs suggest raising it to 0.9 or 0.95 for medical, legal or financial products where any ungrounded claim is a bug. Faithfulness does not catch a grounded answer that draws the wrong conclusion, and it over-penalises products whose answers are meant to go beyond the passages, such as summarisation with commentary.
Answer Precision reads actual_output and expected_output and scores how closely the answer matches the gold answer, combining numeric agreement with tolerance, semantic overlap and factual consistency. The default threshold is 0.8, deliberately higher than most metrics because this is the correctness metric and you want it strict. It is the end-to-end score discussed at the start: the one that tells you whether the product is right, and the one that cannot tell you why it is wrong. One known pitfall is that it penalises paraphrasing, so a semantically equivalent answer can score lower than a verbatim match; when the right answer is a range rather than a fact, the docs point you to G-Eval or Custom Eval with a rubric instead.
Answer Relevancy reads input and actual_output and scores how directly the answer addresses the question, with 1.0 meaning it answers exactly what was asked and 0.0 meaning it talks about something else. The default threshold is 0.6 and it needs no gold answer. On a RAG plan its role is as a companion: an answer that scores well on Answer Precision but poorly on Answer Relevancy is technically correct but wandering, usually because the model summarised the whole retrieved passage instead of answering the question. It is also the generation metric you can run on unlabeled rows, which pairs naturally with Contextual Relevancy for a gold-free baseline.
Note what none of the generation metrics check: whether the retrieved passages were right. Faithfulness will happily score 1.0 on an answer that faithfully reproduces a wrong or outdated passage. That is not a flaw in the metric, it is the separation of concerns doing its job, but it means a high Faithfulness score is a statement about the generator and nothing else. The retriever has its own metrics, and the knowledge base itself has none, which we come back to under what these metrics do not catch.
Reading metrics in pairs
Localization happens when you put a retrieval metric and a generation metric side by side on the same failing rows. The Finding bottlenecks page describes where to look on the run page: the per-metric charts under the KPI tiles show each metric's score distribution with a threshold marker, and a metric with most of its mass below the marker is the one dragging the pass rate down. Once you see which metric is low, the pattern of which others are high tells you the stage.
The first pattern is Contextual Recall low, regardless of what the generation metrics say. Retrieval did not bring back the facts the gold answer needs, so the generator never had a chance. Fix retrieval first; the docs are blunt that no amount of prompt work will fix the answer in this case. If Answer Precision happens to be high on the same rows, the model is answering from its own knowledge, which is a separate concern covered by the next pattern. The second pattern is Contextual Recall high and Answer Precision low. Retrieval did its job and generation dropped the ball: the facts were in the context and the model did not use them correctly. This is prompt or model work, and the fastest way to see what went wrong is the judge's reasoning text in the row expand view under Metrics, which spells out which claim in the expected answer the actual answer missed or contradicted.
The third pattern is Faithfulness low and Answer Precision high. The model is answering right but with ungrounded claims, which is common in RAG products where the target has enough general knowledge to guess correctly even when the retrieved context is wrong. This looks like success in a pass-rate chart and is one of the most dangerous states a RAG system can be in, because the correct answers are coming from the model's memory and will stop being correct the moment the question moves outside what the model memorised. The docs call it bad on trust and say to fix retrieval. The fourth is the mirror image, Answer Precision low and Faithfulness high: the model is grounded in the context but drawing the wrong conclusion from it. Everything it said is in the passages; it just read them wrong. That is a prompt or model issue, not retrieval. A fifth, Contextual Relevancy low and Answer Precision high, is the gold-free version of the third: the model answers correctly despite noisy retrieval, and the docs say to fix retrieval anyway.
A worked example makes the patterns concrete. Suppose you run an internal assistant over your product documentation and one dataset row asks whether the Faithfulness threshold should be raised for a medical product, with an expected output saying yes, to 0.9 or 0.95. In one run the retriever returns the general metrics page, which mentions that the default threshold for Faithfulness is 0.7, and the assistant answers that 0.7 is fine because it is the default. Every claim in that answer is in the passages, so Faithfulness passes, but Answer Precision fails: the fourth pattern, a grounded wrong conclusion. Now suppose instead the retriever returned the RAG metrics page with the compliance advice on it, and the assistant still said 0.7. Contextual Recall passes, Answer Precision fails: the second pattern, retrieval delivered and generation ignored it. And if the retriever returned nothing about thresholds at all and the assistant still answered 0.95 from memory, Contextual Recall fails while Answer Precision passes: the model guessed well, and the first pattern says to fix retrieval before that luck runs out. Three identical-looking correctness failures (or one suspicious success), three different fixes, and the pair of scores is what separates them.
Getting the context onto the row
Every pattern above depends on the retrieval_context column holding the passages the model actually saw, and this is where most RAG evaluations silently go wrong. Faithfulness only sees that column. If your retrieval pipeline passes context some other way, embedded in the system prompt, streamed through a tool call, or appended to the user message by middleware, the column is empty, and Faithfulness scores as if there had been no grounding at all. Every claim in the answer is then an ungrounded claim, the metric fails on every row, and the run looks like a catastrophic hallucination problem when the product is behaving fine. The Metrics hub lists "a metric that always fails" as a sign of wrong columns, and for RAG plans this is the usual cause.
There are two ways to fill the column. The first is on the dataset: you write the passages into retrieval_context when you build the row. This works for a frozen knowledge base and a frozen retriever, and it is useful when you want to evaluate the generator in isolation, because every run gets the same passages and any change in Faithfulness or Answer Precision is a change in generation. The weakness is that it does not evaluate the retriever at all, and if the retriever changes, the dataset is now testing context the model no longer receives. The second way is on the connector: the Retrieval context path field in the connector's output settings points at where the retrieved passages live in your AI system's response, and EvaliQA fills the column at run time from the live retrieval. This is the setup that makes the retrieval metrics meaningful, because Contextual Recall is now scoring what the retriever did on this run, not what it did when you built the dataset. It requires your AI system to return its retrieved passages in the response, which is a small engineering change that pays for itself the first time Contextual Recall drops after a re-index.
Order is the second setup detail. Contextual Precision needs ranked context: if the passages arrive as an unordered set, the metric cannot tell top from bottom and produces noise. Preserve the order the retriever returned, and set the metric's top_k parameter to the number of chunks the model is actually given, so a retriever that returns 20 and a prompt that uses 5 are scored on the 5. For production traces scored by online evaluation, the same idea applies through the trace field mapping: the Online evaluation page describes mapping retrieval_context to a dotted path or JSONPath into the trace, usually into metadata or spans, so Faithfulness on live sessions reads the real passages rather than nothing.
A starting set and thresholds
The temptation on a RAG plan is to attach all six metrics at once. The docs recommend against it, and the reason is specific to localization: Answer Precision, Answer Relevancy, Contextual Recall and Faithfulness on the same plan largely ask the same question several ways, and each LLM-based metric is one judge call per row, so a 300-row dataset with four metrics costs 1,200 judge calls before the target's own 300. Worse, the aggregate pass rate becomes a blend of six opinions and stops pointing anywhere. Start narrow with metrics that read different columns, so that their disagreement means something.
The starting set the docs give for a RAG assistant is Answer Precision, Faithfulness, Contextual Recall and Toxicity. That set covers the two stages and the end result with one metric each: Contextual Recall for whether retrieval delivered the facts, Faithfulness for whether generation stayed inside them, Answer Precision for whether the user got the right answer, and Toxicity as the cheap safety check that belongs on every plan that talks to people. Those three RAG metrics are exactly the ones the pair patterns use, which is not a coincidence. If you have no gold answers, swap Answer Precision and Contextual Recall for Answer Relevancy and Contextual Relevancy, and accept that the patterns you can read are the weaker gold-free ones. Add Contextual Precision only when someone is actively tuning the ranker and will act on the ordering score.
Thresholds turn scores into verdicts, and the defaults are starting points rather than answers. A row passes a metric when its score is at or above the threshold, and a row's overall verdict is the AND of every metric on the plan, so with four metrics a row fails if any one of them does. Read per-metric pass rates rather than the aggregate for that reason. The defaults for the RAG set are 0.8 for Answer Precision, 0.7 for Faithfulness and 0.7 for Contextual Recall, and the docs describe tightening as legitimate: if a metric passes 100 percent of rows on your golden set, it is not discriminating, and raising it from 0.7 to 0.85 and re-running is the right move. For compliance-sensitive products, raise Faithfulness to 0.9 or 0.95, because a single ungrounded claim in a medical or financial answer is an incident rather than a fuzziness problem.
Pin the judge. Every metric here except Restricted Refusal routes through the plan's judge model, and that model is part of the measurement instrument. Rotate it between two runs and the pass rates stop being comparable, which matters most for exactly the pair reading this post is about: a 10-point drop in Contextual Recall between runs on different judges is a judge difference until proven otherwise. Use a cheap model for the system under test if you like, but keep a strong judge and keep it the same across every run you intend to compare.
What these metrics do not catch
The retrieval and generation metrics score the two stages of a RAG system against each other. They do not score the knowledge base. If a passage in your index is wrong or out of date, a retriever that returns it scores well on Contextual Relevancy, a generator that reproduces it scores 1.0 on Faithfulness, and the only metric that will notice is Answer Precision, and only if your gold answer was written from a correct source rather than from the same stale document. A content problem looks like a generation problem under these metrics. The honest answer is that no automated metric catches it; you catch it by reading the judge's reasoning on failed Answer Precision rows and noticing that the answer agrees with the passage and the passage disagrees with reality.
They also do not catch chunking problems directly. A knowledge base chunked so that a table's header is in one chunk and its rows in another produces passages that are individually on topic and collectively useless, and Contextual Relevancy will pass them. Contextual Recall will usually fail, because the gold answer's claims are not supported by any single chunk, but the metric says "retrieval missed facts", not "your chunker split the facts". Retrieval metrics localize the failure to the retrieval stage; what inside that stage is broken, embedding, ranking, chunking, index freshness, is for you to find by reading the retrieved passages on the failing rows.
All six metrics are single-turn. The catalog lists them as applicable to single-turn evaluation plans only, because each reads one input and one output and has no view of a conversation history. A RAG assistant inside a multi-turn product that retrieves on turn 3 based on something the user said on turn 1 cannot be scored for grounding with these metrics as they stand; the Agent metrics cover conversation-level properties, but not retrieval grounding per turn. Latency and cost are also outside the metric set: they are always-on KPI tiles on the run page, and a retriever that doubled its chunk count will show up as a jump in average input tokens before it shows up in any score.
Finally, they are LLM-judge metrics, which means slight non-determinism on every row. Two identical runs can differ by a few percentage points on any of them, which is why the rerun-before-you-fix rule exists. A pair pattern that holds across two unchanged runs is a diagnosis; a pattern that appears in one run and vanishes in the next was the judge. Budget for the second run as part of the method rather than as an afterthought.
If you already have a RAG plan with only Answer Precision on it, the single highest-value change is adding Contextual Recall and Faithfulness, then opening the per-metric charts on the next run and reading the pair on the failing rows. If you do not yet have retrieved passages on the row, fix that first: set the Retrieval context path on your connector so the column fills at run time, or write the passages into the dataset if you want to hold retrieval fixed while you work on the prompt. Then check one failed row's Retrieval context in the row expand to confirm the column holds what you expect before trusting any grounding score.
Once the pair reading points at a stage, run the plan a second time unchanged to make sure the pattern is stable, read the judge's reasoning on five to ten failing rows for the low metric, and take the fix to the team that owns that stage with the pattern name and the rows attached. For the mechanics of the run page, the Evaluation results page covers what each part tells you, and Comparing runs covers how to hold the dataset and the judge constant so the delta after your fix means something.




