On most RAG test plans we see, Answer Relevancy and Faithfulness sit next to each other in the metric list, both are scored by the same LLM judge, both come back as a number between 0 and 1 where higher is better, and both get read as "answer quality". That reading is where the trouble starts. A team sees Faithfulness passing on 90 percent of rows and Answer Relevancy passing on 60 percent and concludes that the system is mostly fine with a few weak answers, when in fact the two numbers describe two different defects in two different parts of the pipeline. Treated as one signal they blur into an average that points nowhere. Treated as two, they tell you which component to open.
The confusion is understandable because the metrics share a vocabulary. Both describe "the answer", both involve the user's question, both are judged by a model reading text, and in EvaliQA both live in the same RAG category of the metrics step. But they read different columns and ask different questions. Answer Relevancy asks whether the response addresses what the user asked. Faithfulness asks whether every claim in the response is supported by the passages the system retrieved. Each one catches a bug the other walks straight past. This post is about those two bugs, with one worked example that fails each, and about what to do when the two scores disagree.
What Answer Relevancy measures
Answer Relevancy reads two columns, input and actual_output, and scores how directly the response addresses the user's question. A score of 1.0 means the response answers exactly what was asked, and 0.0 means it talks about something else. Nothing else enters the judgement. The metric does not see a gold answer, does not see the retrieved passages, and does not know whether the facts in the response are true. It only compares the shape of the question with the shape of the answer and asks whether one fits the other.
That narrowness is the point. Because it needs no expected_output, Answer Relevancy is the metric you can run on any single-turn dataset, including one generated from production traffic where nobody has written reference answers yet. It is the first metric on a brand-new project for exactly that reason. The default threshold in EvaliQA is 0.6, which is deliberately looser than the correctness metrics, because a response can be partly on-topic in a way that a correctness check cannot be partly right. A row that scores 0.55 on relevancy is usually a response that answered a neighbouring question, or buried the actual answer under a paragraph of preamble.
What the judge is really measuring is a kind of alignment between intent and content. Read the judge's reasoning on a handful of low-scoring rows and the explanation is almost always one of three stories. The response answered a more general question than the one asked. The response answered a different specific question that shared keywords with the one asked. Or the response is a refusal or a deflection dressed up as an answer. All three are prompt and routing problems, not knowledge problems, which is why the single-turn test plan guide files "Answer Precision fine, Answer Relevancy bad" under prompt issues.
The metric has a clear boundary. It is built for one input and one output, so it is the wrong shape for multi-turn conversations, where the question the user is really asking may have been set up three turns earlier. It is also the wrong question for classification or structured-output tasks, where "did the response address the question" is trivially yes and the real question is whether the label or the JSON is right. For those, a deterministic check does the job without a judge call.
What Faithfulness measures
Faithfulness reads three columns, input, actual_output and retrieval_context, and scores whether every claim in the response is supported by the retrieved passages. A score of 1.0 means every claim is grounded. Anything less means the judge found at least one statement that the context does not support, either because the model made it up or because it brought it in from its own training rather than from the passages it was given. This is the RAG-specific hallucination metric. It is not asking whether the answer is true in the world, it is asking whether the answer is true according to the documents the system was told to use.
That distinction matters more than it first appears. A RAG system is a promise that answers come from a controlled corpus, so that when the corpus changes the answers change with it, and so that a reviewer can trace any statement back to a source. A response that is factually correct but not supported by the retrieved passages breaks that promise just as badly as a false one, because it means the model is answering from memory and the corpus is decorative. Faithfulness is the metric that tells you whether the promise holds, which is why it is the one metric the docs put on every RAG plan without exception.
The default threshold is 0.7, and the RAG metrics guide recommends raising it to 0.9 or 0.95 for medical, legal and financial products, where a single ungrounded claim is a bug rather than a quality nuance. The single-turn guide goes further and treats a Faithfulness pass rate below 95 percent on a RAG system as a release blocker. The reasoning is simple. The metric scores claims, and a response with five grounded claims and one invented one is a response that will eventually be quoted back to you by a customer with the invented part highlighted.
Faithfulness has two limits worth stating up front. It only sees the retrieval_context column, so if your pipeline passes retrieved passages some other way, for example by embedding them in the system prompt, the metric sees no grounding at all and scores the response as if it had made everything up. And it over-penalises any product that is meant to go beyond the passages, such as a summariser that adds commentary or an assistant that is allowed to draw on general knowledge. On those, every reasonable addition reads as a fabrication, and the score tells you about the design, not the bug.
The bug Answer Relevancy catches
Take a support assistant that answers questions over a retailer's returns policy. The user asks "Can I return a laptop after 30 days if it is still unopened?". Retrieval brings back a passage that says unopened items may be returned within 45 days of delivery for a full refund, and that opened electronics returned within 30 days are subject to a 15 percent restocking fee. The response the system produces is "Opened electronics are subject to a 15 percent restocking fee when returned within 30 days of delivery." Every word of that is in the passage. Faithfulness has nothing to object to, and if you were only reading Faithfulness this row would look like a success.
It is not. The user asked about an unopened laptop after 30 days and got a sentence about opened electronics within 30 days. The response is accurate, grounded, well written and useless, and the user has to either ask again or guess. Answer Relevancy is the metric that flags this row, because the judge compares the question to the response and sees that the response addresses a neighbouring case rather than the one asked. The judge's reasoning will say something close to "the response discusses opened items and a restocking fee, while the question concerns an unopened item beyond 30 days".
This failure mode is common in RAG systems for a mechanical reason. The retriever surfaces a chunk that contains the right policy, the generator latches onto the most specific sentence in that chunk, and the specific sentence is about a different branch of the policy. Nothing was invented, so no grounding check fires. The system also looks fine on a correctness metric if your expected output happens to mention the restocking fee. The only metric that notices is the one that reads the question and the answer and nothing else. That is why Answer Relevancy belongs on a RAG plan even though it is not a RAG-specific metric.
When this pattern shows up across many rows, the fix is almost never in retrieval. The bottleneck guide describes the signature as Answer Precision fine, Answer Relevancy bad, meaning the model is technically right but talking around the question, and it points at the prompt. In practice that means instructing the generator to identify the user's specific case before answering, or restructuring the corpus so that a chunk does not pack several policy branches into one paragraph. Either way, the diagnosis came from one metric disagreeing with the others, not from the aggregate pass rate.
The bug Faithfulness catches
Same assistant, same question, same retrieved passage. This time the response is "Yes, unopened laptops can be returned within 60 days for a full refund, and we also cover the return shipping." Read it as the user would and it is a model answer. It addresses the exact question, it is confident, it gives a number and a bonus detail. Answer Relevancy scores it high, because the response is precisely about what was asked. Any reader who had not seen the passage would take it as the policy.
The passage says 45 days, not 60, and says nothing at all about shipping. Two of the three claims in the response are unsupported. Faithfulness is the metric that fails this row, and the judge's reasoning will list the claims it checked and mark the 60-day window and the shipping promise as absent from the context. This is the hallucination that matters in production, because it is the one that looks finished. A nonsense answer gets caught by the user. A plausible answer with a wrong number gets acted on, and the retailer finds out when a customer shows up on day 50 with a screenshot.
Where did the 60 days come from? Usually from somewhere real. The model may have seen a different retailer's policy in training, or an older version of this one, or the number may simply be the most common value for that sentence shape. The retrieval was correct, the model had the right passage in front of it, and it still answered from memory. The bottleneck guide calls this pattern Faithfulness low, Answer Precision fine and notes that it is common in RAG products where the model has enough general knowledge to produce a good-looking answer regardless of what retrieval returned. It is bad on trust precisely because the system often gets away with it.
No metric that reads only the question and the answer can see this failure. Answer Relevancy cannot, because the response is relevant. Answer Precision cannot unless your expected output happens to disagree on the exact detail that was invented, and even then it reports a mismatch rather than a grounding failure. Only a judge that is handed the retrieved passages and asked to check each claim against them can separate "the model answered from the corpus" from "the model answered from memory and the corpus agreed by luck". That is the whole job of Faithfulness, and it is the reason the metric requires the retrieval_context column rather than treating it as optional.
Why the two get confused
The names do most of the damage. Both metrics are about the answer, both involve relevance in some everyday sense, and in the EvaliQA metrics step both are listed in the RAG category, which the docs admit is a slightly misleading label because it also holds the general-purpose single-response metrics. A reader skimming the list sees two LLM judge metrics with 0.6 and 0.7 thresholds and reasonably assumes they are two flavours of the same check. They are not, and the cleanest way to see it is to look at what each one is given to read.

The second source of confusion is that both metrics tend to move together on healthy systems. On a well-behaved RAG assistant, most rows are both relevant and grounded, so the two pass rates track each other and it is easy to start thinking of them as one measurement with two readouts. The value of running both only shows on the rows where they split. The hub page for metrics warns against stacking Answer Precision, Answer Relevancy, Contextual Recall and Faithfulness on one plan because they overlap, and that warning is right about the overlap in general. But Answer Relevancy and Faithfulness specifically are the pair that overlap least, because one reads the context and the other cannot.
The third source is the judge itself. Both metrics use the plan's judge model, both produce a short reasoning trace, and if you only ever read the score you never see that the two traces are about different things. The fastest cure for the confusion is to open the row expand view on a failing row and read the reasoning under each Metrics card side by side. One will talk about the question. The other will talk about claims and passages. After doing that on five rows the two metrics stop looking alike.
Reading the two scores together
Once you treat them as separate instruments, every row falls into one of four cells, and each cell has a different meaning. The two agreeing cells are the easy ones. Both high means the response answered the question using the passages, which is what a RAG system is for. Both low means the response is off-topic and ungrounded at once, which usually indicates that retrieval returned nothing useful and the model filled the gap with a generic answer to a nearby question. On that cell, check the context chunks in the row expand view before blaming the generator. If the chunks are empty or irrelevant, the problem is upstream and a Contextual Relevancy or Contextual Recall metric will confirm it.
The two disagreeing cells are where the diagnosis lives. Relevancy high and Faithfulness low is the invented 60-day window from the example above. The model understood the question and answered it from memory, so the fix is in generation, typically a stronger instruction to answer only from the provided passages, a stronger model, or a lower temperature on the target. Relevancy low and Faithfulness high is the restocking-fee answer. The model stayed inside the passages but picked the wrong part of them, so the fix is in how the prompt frames the user's case or in how the corpus is chunked. Two different failures, two different owners, and the aggregate pass rate would have shown the same number for both.
Answer Relevancy | Faithfulness | What happened | Where to look |
|---|---|---|---|
High | High | Answered the question from the passages | Nothing, this is the goal |
High | Low | Answered the question from memory | Generation prompt, target model, or an unpopulated context column |
Low | High | Quoted the passages but answered a different question | Prompt framing, chunk boundaries in the corpus |
Low | Low | Off-topic and ungrounded | Retrieval first, then generation |
One habit makes this table usable. Before you form a diagnosis from a split, run the plan again unchanged. Both metrics are LLM judge metrics, and the same judge scoring the same row twice can land on slightly different numbers, so a row sitting at 0.68 on a 0.7 threshold can flip between runs for no reason connected to your system. If the split persists across two identical runs it is real. If it moves, lower the judge temperature or pin a stronger judge before you spend a week on the prompt. The bottleneck guide calls this the rerun before you fix rule and it applies with extra force to any analysis that depends on two metrics disagreeing.
What neither metric catches
Running both metrics does not give you a complete picture of a RAG system, and it is worth being precise about the gaps. The first is correctness. Neither metric compares the response to a gold answer. A response can be on-topic and fully grounded and still wrong, because the retrieved passage itself was out of date or because the model drew the wrong conclusion from a correct passage. The bottleneck guide describes that second case as Answer Precision low, Faithfulness fine, and it is invisible to the two metrics in this post. If you have expected outputs, Answer Precision is the metric that fills this gap, and the docs recommend it as the primary metric on any supervised dataset.
The second gap is retrieval. Faithfulness checks the answer against whatever was retrieved, which means a perfectly faithful answer to a badly retrieved passage scores 1.0. If the retriever brought back the shipping policy instead of the returns policy and the model dutifully summarised the shipping policy, Faithfulness is satisfied, and Answer Relevancy may or may not fire depending on how far off the topic drifted. Contextual Recall, which needs an expected output, tells you whether retrieval brought back enough to answer at all, and Contextual Precision tells you whether the useful chunks were ranked first. The RAG metrics guide pairs Faithfulness with those two precisely so that you can separate "retrieval was bad" from "retrieval was fine and generation drifted".
The third gap is the one that produces the most confusing runs, and it is a plumbing problem rather than a quality problem. Faithfulness only reads the retrieval_context column. If that column is empty, because your connector returns the passages in a field EvaliQA was never told about, or because your dataset was generated without context, the metric does not error out. It scores every response as ungrounded, because from its point of view there was nothing to ground on. A Faithfulness pass rate near zero while Answer Relevancy and Answer Precision look healthy is almost always this. Open a row, look at the Context chunks block on the left of the row expand view, and if it is empty the metric is telling you about your configuration, not your model.
The last gap is scope. Both metrics are single-turn. They read one input and one output, so on a multi-turn plan they can score an individual turn but they cannot tell you whether the conversation as a whole stayed on task or stayed grounded across a user's follow-ups. The agent metrics that read the turns column exist for that. And neither metric says anything about safety, tone, format or cost. A response can be relevant, faithful and toxic. The docs suggest keeping Toxicity on every plan for that reason, and it is cheap enough that there is no argument against it.
Setting them up in EvaliQA
Both metrics attach on the test plan, and the wizard already filters them to single-turn evaluation plans, so you will not see them offered on a multi-turn or red teaming plan. The one step that trips people up is the last one, making sure the retrieved passages actually reach the retrieval_context column at run time. The sequence below is the one we use on a new RAG project.

Two further details save time later. Clicking the metric name on a Metrics card opens the verbose log with the full judge prompt and response, which is the quickest way to see exactly which claims Faithfulness checked and which it rejected. And if a run shows 100 percent pass on either metric across your whole golden set, the threshold is probably too low to discriminate on your data. The hub guide's advice is to tighten it, for example from 0.7 to 0.85, and run again. A metric that always passes costs a judge call per row and tells you nothing.




