Evaluating RAG Retrieval Before Changing the Language Model

Measure RAG evidence with Recall@k, MRR, and nDCG, using tested Python code to expose missing passages, ranking tradeoffs, and context-packing failures.

12 min read

A support assistant finds the general refund policy, quotes it accurately, and still gives the wrong answer. The customer bought a product covered by an exception in a different document. The first result was relevant. The answer needed two pieces of evidence.

That constructed example exposes a common weakness in RAG evaluation: a successful search hit gets mistaken for a complete evidence set. Switching the language model may change the wording without repairing the missing input. Before changing the generator, inspect what the retrieval pipeline actually delivered.

This article builds a small, dependency-free Python evaluator for Recall@k, reciprocal rank, and nDCG@k. The rankings are hand-authored so their differences are easy to audit. They are a metric laboratory, not a benchmark of an embedding model, vector database, or production RAG system.

Put the measurement at the evidence boundary

The original Retrieval-Augmented Generation paper combines retrieval from an external index with a language generator. For application debugging, that separation gives us a useful boundary: the evidence supplied to generation can be inspected independently of the answer.

In a pipeline with candidate retrieval, reranking, and context packing, save the output at each stage. A relevant passage can be absent from the candidate set, pushed below the reranking cutoff, or removed while fitting the prompt into a token budget. Those are three different defects with different owners.

Record ordered passage IDs, source revisions, scores, filters, and the final context text. A passage ID alone is insufficient if a formatter can truncate the sentence containing the answer. Include the query’s permission scope and effective date: evidence from another tenant or an obsolete policy is not acceptable evidence for that request.

For one failing query, keep the prompt and generator settings fixed and replace retrieved context with a manually verified evidence set. If the answer improves, you have evidence that the original context contributed to the failure. If it does not, inspect instruction following, reasoning, answer grading, and whether the supposedly sufficient context really answers the question. This is a controlled diagnostic intervention, not proof that a particular retrieval algorithm is best.

Decide what a relevant item means

A judgment file, often called qrels, maps each query and item ID to a relevance grade. Choose the item before assigning grades. For the example below, each ID represents one independently retrievable passage:

  • 0: irrelevant to the information need.
  • 1: useful background, but weak support for the requested answer.
  • 2: substantial support for part of the answer.
  • 3: direct support for a central fact the answer requires.

Here, every grade above zero counts as relevant for binary metrics. That choice deliberately lets background satisfy reciprocal rank. A stricter application could use a threshold of two, but it must change both the hit test and the recall denominator. Record the threshold beside the score; it is part of the experiment.

Document-level and passage-level judgments answer different questions. Mapping ten overlapping chunks back to one parent document can measure source discovery. It cannot establish that the exact answer-bearing span reached the prompt. Conversely, counting every overlapping chunk as separate relevant evidence can reward repetition. Define a stable evaluation unit, then keep it fixed across compared runs.

For a large corpus, exhaustive judgments are rarely practical. Pooling candidates from several retrieval systems is a standard way to select material for assessment. Include alternatives to the current system so it does not define its own answer key. Stanford’s discussion of relevance assessment explains this approach. Unjudged material remains unknown; it is not evidence of irrelevance.

Three metrics, three different questions

Recall@k asks how many known relevant items appeared within the first k results, divided by the total number of known relevant items for the query. Missing items stay in that denominator. This follows the standard definition of recall, applied to a ranked prefix.

RR@k is the reciprocal of the first relevant rank within that prefix, or zero when no hit appears. The mean over queries is MRR@k. NIST’s trec_eval reciprocal-rank implementation illustrates the first-hit definition; this example additionally applies an explicit cutoff. Once rank one is relevant, RR is already perfect, even if other necessary evidence is missing.

nDCG@k rewards higher relevance grades near the top. Here the gain is 2**grade - 1 and the discount is log2(rank + 1). Divide the resulting DCG by the best possible DCG at the same cutoff, using all available judgments. This is the exponential-gain formulation described in Stanford’s ranked-retrieval evaluation chapter. State the gain formula when comparing tools: a linear-gain implementation need not produce the same values.

None of these metrics directly tests whether an answer is correct, well cited, or supported in full. They summarize rankings under a particular judgment policy.

A small evaluator that rejects ambiguous inputs

Save the following as retrieval_metrics.py. It uses only the standard library and was tested with Python 3.13.4.

from math import log2
from statistics import fmean


def score_query(qrels: dict[str, int], ranked: list[str], k: int) -> dict[str, float]:
    """Score a fully judged, answerable query; positive grades are relevant."""
    if type(k) is not int or k <= 0:
        raise ValueError("k must be a positive integer")
    if any(not isinstance(doc, str) or not doc for doc in qrels):
        raise ValueError("judgment IDs must be nonempty strings")
    if any(type(grade) is not int or not 0 <= grade <= 3
           for grade in qrels.values()):
        raise ValueError("grades must be integers from 0 to 3")
    if any(not isinstance(doc, str) or not doc for doc in ranked):
        raise ValueError("ranked IDs must be nonempty strings")
    if len(ranked) != len(set(ranked)):
        raise ValueError("duplicate ranked IDs")
    if set(ranked) - qrels.keys():
        raise ValueError("all returned IDs must have judgments")

    relevant = sum(grade > 0 for grade in qrels.values())
    if relevant == 0:
        raise ValueError("no positive judgments; evaluate abstention separately")

    grades = [qrels[doc] for doc in ranked[:k]]
    recall = sum(grade > 0 for grade in grades) / relevant
    rr = next((1.0 / rank for rank, grade in enumerate(grades, 1)
               if grade > 0), 0.0)

    def dcg(values: list[int]) -> float:
        return sum((2 ** grade - 1) / log2(rank + 1)
                   for rank, grade in enumerate(values, 1))

    ideal = sorted(qrels.values(), reverse=True)[:k]
    return {"recall": recall, "rr": rr, "ndcg": dcg(grades) / dcg(ideal)}


def evaluate(qrels: dict[str, dict[str, int]], run: dict[str, list[str]],
             k: int) -> tuple[dict[str, dict[str, float]], dict[str, float]]:
    """Macro-average over exactly the expected query set."""
    if not qrels or qrels.keys() != run.keys():
        raise ValueError("run must contain exactly the nonempty query set")
    per_query = {qid: score_query(labels, run[qid], k)
                 for qid, labels in qrels.items()}
    means = {metric: fmean(row[metric] for row in per_query.values())
             for metric in ("recall", "rr", "ndcg")}
    return per_query, means

The evaluator requires judgments for every returned ID, including IDs beyond the scoring cutoff. This strict contract suits the tiny, fully judged fixture. On a real evaluation pool, either complete those judgments or use an explicitly documented policy for incomplete labels. Silently assigning unknown items a zero can penalize a system for discovering useful material the assessors have not seen.

Duplicate IDs are rejected rather than counted twice. If production performs deduplication, evaluate the actual resulting list and record where deduplication happened. Do not remove duplicates and pull extra candidates only during evaluation: that quietly changes the retrieval budget.

A query with no positive judgments raises an exception. Recall has no positive denominator in that case. Treating it as a perfect result would inflate the average; treating it as an ordinary failed retrieval would mix two different tasks. We will handle unanswerable requests separately.

Run two rankings through the same judgments

Save this second file as demo.py beside the evaluator. The names stand for hypothetical passages about refunds, request limits, and recovery. They are labels for the experiment, not claims about a real product’s policies.

from retrieval_metrics import evaluate

# Hand-authored judgments and rankings, not a model benchmark.
QRELS = {
    "refund": {"policy": 3, "exception": 2, "overview": 1, "noise": 0},
    "limits": {"limit": 3, "scope": 2, "noise": 0},
    "restore": {"restore": 3, "snapshot": 2, "noise": 0},
}
RUNS = {
    "A": {
        "refund": ["overview", "policy", "exception"],
        "limits": ["limit", "noise"],
        "restore": ["noise", "restore", "snapshot"],
    },
    "B": {
        "refund": ["policy", "exception", "overview"],
        "limits": ["noise", "scope", "limit"],
        "restore": ["restore", "noise", "snapshot"],
    },
}

if __name__ == "__main__":
    print("run query     Recall@3 RR@3/MRR@3 nDCG@3")
    for name, run in RUNS.items():
        rows, means = evaluate(QRELS, run, k=3)
        for qid, scores in [*rows.items(), ("MEAN", means)]:
            print(f"{name:3} {qid:9} {scores['recall']:.4f}   "
                  f"{scores['rr']:.4f}      {scores['ndcg']:.4f}")

Run python3 demo.py. These are the actual results; the MEAN rows use MRR@3, while individual query rows show RR@3.

run query     Recall@3 RR@3/MRR@3 nDCG@3
A   refund    1.0000   1.0000      0.7364
A   limits    0.5000   1.0000      0.7872
A   restore   1.0000   0.5000      0.6653
A   MEAN      0.8333   0.8333      0.7296
B   refund    1.0000   1.0000      1.0000
B   limits    1.0000   0.5000      0.6064
B   restore   1.0000   1.0000      0.9558
B   MEAN      1.0000   0.8333      0.8541

For refund, both runs retrieve all three useful passages and put a relevant item first. Recall and RR cannot distinguish them. Run A starts with grade-one background; B starts with direct evidence and follows with substantial support. nDCG moves from 0.7364 to 1.0000.

For limits, A finds the strongest passage immediately but misses scope. Its RR is 1.0000 and recall is only 0.5000. B recovers both passages, but an irrelevant result comes first and the strongest evidence arrives third. Recall improves while RR and nDCG fall. If the answer requires both the limit and its scope, that lower-ranked result may still be operationally preferable. The metric alone cannot choose the product’s tradeoff.

Across all three queries, MRR stays at 0.8333. Mean recall rises from 0.8333 to 1.0000, and mean nDCG rises from 0.7296 to 0.8541. An MRR-only dashboard would report no change. An aggregate nDCG dashboard would show improvement while hiding the regression on limits. Retain per-query rows.

These three examples establish how the calculations behave. They do not establish statistical significance, expected business impact, or a reason to deploy B.

Keep missing evidence in the denominator

The easiest implementation mistake is to build the ideal ranking from retrieved items alone. A system that returns just one useful passage could then receive nDCG of one, despite missing a better passage entirely. The ideal ranking must come from the judgment set, independently of the run being scored.

Short result lists need the same discipline. The limits ranking in A returns two items for k=3. It receives no fictional third hit, and neither denominator shrinks to reward the short list. Empty results on an answerable query produce zeros.

The aggregate is a macro-average: each query receives equal weight. It is not a single pooled count of hits divided by all relevant items across queries. Pooling would give more influence to questions with many labeled passages. Also, a missing query is a data error, while a query present with an empty list is a retrieval outcome. The evaluator refuses to quietly drop failed requests from the average.

These executable checks cover several of those boundaries. Save them as test_article.py and run python3 -m unittest test_article -v.

import math
import unittest
from retrieval_metrics import score_query


class MetricChecks(unittest.TestCase):
    def test_relevant_result_at_rank_two(self):
        s = score_query({"a": 3, "x": 0}, ["x", "a"], 2)
        self.assertEqual(s["recall"], 1.0)
        self.assertEqual(s["rr"], 0.5)
        self.assertAlmostEqual(s["ndcg"], 1 / math.log2(3))

    def test_missing_evidence_stays_in_denominator(self):
        s = score_query({"a": 3, "b": 2}, ["a"], 3)
        self.assertEqual(s["recall"], 0.5)
        self.assertLess(s["ndcg"], 1.0)

    def test_cutoff_can_remove_the_only_hit(self):
        s = score_query({"a": 3, "x": 0}, ["x", "a"], 1)
        self.assertEqual(s, {"recall": 0, "rr": 0, "ndcg": 0})

    def test_bad_data_is_not_a_zero_score(self):
        for ranked in (["a", "a"], ["unjudged"]):
            with self.assertRaises(ValueError):
                score_query({"a": 3}, ranked, 3)


if __name__ == "__main__":
    unittest.main()

The article’s wider validation also checks invalid grades and cutoffs, missing query IDs, macro-averaging, and all 24 orderings of a four-item fixture at five cutoffs. That last check verifies score bounds and nondecreasing recall as the prefix grows. It does not assume nDCG must increase with k: its ideal denominator changes too.

Measure the context that survives packing

A candidate retriever can have excellent recall while the generator receives poor evidence. Measure candidate recall at its configured budget, inspect reranked results at the downstream cutoff, and then assess the passages or spans that survive context packing. Keep these reports separate; k=50 candidates and a final 2,000-token context are different constraints.

The token budget must use the generator’s actual tokenizer and reserve space for instructions, the user request, and the answer. Include separators and metadata in the accounting. If long passages are truncated, reassess the retained spans rather than inheriting the full passage’s grade automatically.

For questions requiring several independent facts, add a task-specific evidence-coverage check. Label required facts and the spans that support each one. Then ask whether the final context contains support for every required fact. Call this required-fact coverage, document its denominator, and keep it separate from standard recall. Ten passages repeating one fact do not cover a second fact.

Alternative passages may support the same fact, so requiring every relevant document can be unnecessarily strict. Contradictory versions also need attention: retrieving both a current rule and a superseded rule can increase document recall without creating a usable answer. A relevance score is only as sound as the time, access, and evidence rules used to assign it.

Turn the laboratory into a release comparison

Build an evaluation set from the requests the application needs to handle. Include exact identifiers, paraphrases, ambiguous terms, multi-document questions, and recent changes. Keep related paraphrases and near-duplicate source families together when splitting development and held-out data, so a tuning loop does not masquerade as generalization.

The BEIR benchmark evaluates retrieval across heterogeneous tasks and domains rather than assuming one narrow collection represents everything. For an application, the practical implication is to inspect meaningful query slices: an average can hide a failure on the very requests that justify the system.

Freeze the query set, judgments, corpus snapshot, chunking rules, permissions, and scoring policy for each comparison. Record model and index configuration, candidate counts, latency, and context size. Change one component at a time when diagnosing causes. If chunking changes item IDs, remap judgments to preserved evidence spans and review the mapping before comparing scores.

Calculate paired differences on the same queries and inspect the largest regressions. With an adequately sized set, uncertainty estimates can help distinguish noise from a dependable improvement; resample whole queries rather than treating passages from one query as independent observations. Three hand-authored questions are useful tests of code, not an adequate release sample.

Keep unanswerable requests in a separate evaluation slice. Measure whether the full system abstains when the authorized corpus lacks sufficient evidence, and whether it wrongly abstains on answerable requests. Returning no documents is not, by itself, successful abstention; the generated response still needs to behave correctly.

Finally, evaluate answer correctness and citation support using the exact final context. A retrieval improvement earns its place when it helps the application under acceptable latency and cost. When verified evidence reaches the prompt and answers still fail, the generator becomes a much clearer target for investigation.

What do you think?

Add your perspective.

Your email address will not be published. Required fields are marked *

What is on your mind?

START WITH A TOPIC
SAVED FOR A QUIET MOMENT

My reading list

Your list is stored only in this browser.

See you in the next story.

New ideas, new stories. The same curiosity.

Open the RSS feed