RAG Evaluation
RAG evaluation means scoring retrieval quality and generation quality separately, rather than judging a RAG system by its final answer alone. A RAG pipeline can fail at either stage independently — retrieval can find the wrong material, or generation can produce a poor answer despite retrieval finding exactly the right thing — and a single pass-or-fail check on the final answer can't tell you which one actually broke.
Why the two have to be scored apart
If the final answer is wrong, there are two very different possible causes. Retrieval might never have found the passage that actually answers the question — a retrieval failure. Or retrieval might have found exactly the right passage, and the model still produced a poor answer from it — a generation failure. Fixing the wrong one wastes effort: tuning chunking and embeddings does nothing for a generation problem, and a more capable model does nothing for a retrieval problem. Scoring the two stages separately is what tells you which fix actually applies.
Scoring retrieval
Retrieval quality asks whether the right material was actually found, independent of what the model did with it. Precision asks: of the chunks retrieved, how many were actually relevant? Recall asks: of the chunks that would have actually helped answer the question, how many did retrieval find? A retrieval step can have high precision and low recall — everything it returns is relevant, but it missed other passages that also mattered — or the reverse, and the two failures call for different fixes.
Scoring generation
Generation quality asks whether the model actually used what it was given well. Faithfulness (sometimes called groundedness) checks whether the answer is actually supported by the retrieved material, rather than adding claims the source never stated. Answer relevance checks whether the answer actually addresses the question that was asked, separately from whether it's faithful to the source — an answer can be perfectly faithful to the retrieved text and still fail to actually answer what was asked.
How this actually gets measured
Retrieval metrics can often be checked against a known-correct set of question-and-relevant-chunk pairs, built ahead of time. Generation metrics like faithfulness and relevance usually can't be checked by exact match, since there's no single correct string for "a faithful answer" — this is one of the more common real uses of LLM-as-judge, scoring whether an answer's claims are actually backed by the retrieved text. The same limitations apply here as anywhere else a judge model is used: it needs periodic calibration against real human review, not blind trust.
A practical example
A support assistant answers a policy question confidently, and the answer turns out to be wrong. Checking retrieval first: did the actual current policy text get retrieved at all? If not, that's a retrieval failure — the fix is in chunking, embeddings, or the query itself. If the right policy text was retrieved and the model still got it wrong — contradicted it, or answered a related-but-different question — that's a generation failure, and the fix is in the instructions or the model, not the retrieval pipeline.
In this guide
FAQ
If a RAG system's final answers are usually correct, does retrieval quality still need to be measured separately?
Yes. A system can look fine on average while still failing regularly on a specific type of question, and only measuring the final answer hides which stage is actually responsible when something does go wrong. Separate retrieval scoring is what lets you catch and localize that failure instead of just knowing something, somewhere, went wrong.
Is a faithful answer automatically a correct one?
No. Faithfulness only checks that the answer doesn't say more than the retrieved material supports — if the retrieved material itself was wrong, outdated, or incomplete, a perfectly faithful answer can still be substantively wrong. Faithfulness measures whether generation stayed within its source, not whether the source itself was right.