RAG evaluation: how to know your answers are right
How to evaluate a RAG system: measure retrieval and answers separately, build a test set from real questions, use model judges with care, and test every change.
By Ali HassanSenior engineer for SaaS and AI products
RAG evaluation means checking two things on their own: whether retrieval pulled back the passages that hold the answer, and whether the model's answer stayed faithful to those passages. If you only ever look at the final answer, you cannot tell a retrieval problem from a generation problem, and you end up guessing at fixes. So the first move is to split the pipeline, score each half against a fixed set of questions whose right answers you already know, and re-run that set every time you change something.
RAG evaluation starts by splitting the problem
A RAG system has two jobs that fail in different ways. Retrieval finds the passages that should answer the question. Generation turns those passages into an answer. A wrong answer can come from either: the right passage never came back, or it came back and the model ignored it or embellished it.
If you score only the end result, those two causes look identical, and you cannot know whether to touch your chunking or your prompt. So measure them separately. I cover how the pieces fit together in RAG architecture; this article is about telling whether that architecture is working.
What retrieval measures
Retrieval is the easier half to measure, because the question is concrete: did the passage that holds the answer come back, and did it come back near the top?
- Recall at k. For each question you know which source should answer it. Recall at k asks whether that source is in the top k results. If it is not, nothing the model does afterwards can save the answer.
- Precision. Of the passages that came back, how many were actually relevant? Low precision floods the prompt with noise and pushes the real passage out of the model's attention.
- Ranking. Was the best passage first, or buried at position nine? A system can have good recall and still answer badly if the useful passage sits below three irrelevant ones.
Retrieval metrics are cheap to compute and do not need a model to judge them. You can run them on a laptop, which makes them the first thing to check when answers go wrong.
What generation measures
Generation is harder, because "is this answer good" is a judgement, not a lookup. Break it into parts you can score one at a time:
- Groundedness, or faithfulness. Does every claim in the answer trace back to a retrieved passage? An answer can be fluent, confident and entirely invented. Groundedness is the single most important generation measure, because an ungrounded answer in a support or review tool is worse than no answer.
- Answer relevance. Does the answer address the question that was asked, rather than a nearby one? You score this against the question, not the sources.
- Citation correctness. When the answer cites a source, does that source actually say what the answer attributes to it? A citation that points at the wrong passage looks trustworthy and is not.
- Correct refusal. When the knowledge base holds no answer, does the system decline, or does it guess? Refusal is a measurable behaviour: feed it questions you know are not covered and check that it says so.
One table to keep the measures straight
| What to measure | Question it answers | How to measure |
|---|---|---|
| Recall at k | Did the right passage come back in the top k results? | Check whether the expected source appears in the retrieved set for each test question. |
| Precision | How much of what came back was relevant? | Count relevant passages against the total retrieved. |
| Ranking | Was the best passage near the top, not buried? | Look at the position of the first relevant hit across the test set. |
| Groundedness / faithfulness | Does every claim trace to a retrieved passage? | Judge each claim against the sources and flag anything unsupported. |
| Answer relevance | Does the answer address the question asked? | Score the answer against the question, not against the sources. |
| Citation correctness | Do the cited sources say what the answer claims? | Match each citation back to its passage and check it supports the claim. |
| Correct refusal | Does it decline when the answer is not in the knowledge base? | Feed questions with no supporting source and check it refuses rather than inventing. |
Build a test set from real questions
None of these measures mean anything without a fixed set of questions to run them on. Build it from questions people actually ask, not ones you invent at your desk. Pull them from support tickets, search logs, or whatever your users already send.
For each question, record two things: the answer a knowledgeable person would give, and the source or sources that answer should come from. The expected source is what lets you score retrieval without a human in the loop every time.
Include the awkward cases on purpose. Questions the knowledge base does not cover, so you can measure refusal. Questions with terms and product names that paraphrase-only search misses. Questions in other languages if your users write in them. A test set that holds only easy questions will tell you everything is fine while real traffic quietly fails.
Start small. A hundred good questions with known sources is more useful than a thousand you never curated.
Using a model as judge
Groundedness, relevance and refusal are tedious for a person to score by hand on every run, so a common approach is to use a model as the judge: give it the question, the retrieved passages and the answer, and ask it to rate whether the answer is supported.
A model judge is good at the mechanical parts of this. It is patient, it is consistent across a thousand cases in a way a tired reviewer is not, and it can point at the sentence it thinks is unsupported. For groundedness and citation checks, where the task is "does this text support that claim", it does reasonably well.
It also has biases worth knowing. Judges tend to prefer longer, more confident answers, and to prefer answers written in the same style as their own. They can be swayed by fluent writing over correct content. And a judge built on the same model family that wrote the answer may grade its own habits too kindly.
The guard against this is simple: check the judge against human labels. Have a person score a sample of the same cases, compare the two, and see where they disagree. If the judge and the humans line up, you can trust the judge to run at scale. If they do not, fix the judge's prompt, or keep more of the scoring with people, before you rely on its numbers.
Human review still has a job
Automated scores tell you what changed. They do not always tell you whether the change matters. Keep a person reading a sample of answers, especially the ones the judge marks as borderline and the ones real users complained about.
Human review is where you catch the failures no metric was written for: an answer that is grounded but unhelpful, a tone wrong for the audience, a refusal that should have been an answer. Use it to calibrate the judge, not to replace it.
Run the test set on every change
Here is the point of all this work. A RAG pipeline has many moving parts, and a change to any one of them can move accuracy in either direction: the prompt, the model, the chunk size, the embedding model, the number of passages retrieved. Each of these looks like an improvement in isolation and can quietly make answers worse.
So treat the test set as a regression check. Run it before and after every change and compare the scores. A new embedding model that lifts recall but drops refusal accuracy is not an upgrade until you have seen both numbers. Without the test set, the first you hear of a regression is a customer.
On Evoriqa, my own support platform, test cases can be run against the live agent and scored by a model judge, and individual replies can be scored in the background. Each answer stores what it was based on, the chunks, their scores and the prompt, so a team can see why it said what it said rather than arguing about it. On a financial-statement review tool I built for a firm, each model finding has to include a short verbatim quote that is matched back to the document to give it a section and page number, so a reviewer can check the citation instead of trusting it.
Refusal belongs in the regression set too. Evoriqa gives a fixed reply when its knowledge base holds no good answer, instead of guessing, and that behaviour is exactly the kind that a prompt change can silently break.
Watch what production tells you
A test set is a snapshot. Real traffic keeps moving, and it hands you evaluation signals for free if you collect them. These sit alongside your observability and often matter more than any offline score, because they come from the people the answers are for.
The useful signals are the ones that show a human reacting to a drafted answer:
- Edits. When a person rewrites a drafted reply before sending it, the size of the edit is a measure of how wrong the draft was. Small tidy-ups mean the system is close; heavy rewrites mean it is not.
- Approvals. How often a draft is sent as-is tells you how often it was right enough to trust.
- Escalations. When a conversation gets handed to a person, something was not answerable or not answered well.
Evoriqa's shadow mode is built around this. The agent drafts every reply into the team inbox, and nothing reaches the customer until a person sends it. Per-channel numbers, drafts, approvals and how much people edit, show when a channel is ready to run on its own. That is evaluation on live traffic with a safety net, and it is the most honest score you can get.
When to bring in help
If your RAG answers cannot be trusted yet and you are not sure whether the fault is retrieval or generation, that is a measurement problem before it is a fixing problem. I run a short LLM cost and RAG quality audit that measures where answers are ungrounded, where retrieval misses, and which change moves accuracy most, with the findings ranked by cost and risk. A good place to start if you have an AI feature that demos well and cannot yet prove it is right.


