Home / Articles / Six RAG evaluation metrics that matter in production

This article is published in English.

Six RAG evaluation metrics that matter in production

Recall@K, nDCG, MRR, faithfulness, latency, and cost—how to debug retrieval that looks better on paper while answers get worse.

2244 words

Early RAG tutorials make the path look short: split documents, embed them, store vectors, pick a small top_k, call the model. That recipe can look healthy in a demo. Production is less kind.

Questions arrive messy. Source documents contradict each other. Some prompts need evidence from several places at once. Users invent phrasings that never appeared in the original test set. A setup that scored well on a canned benchmark starts returning answers that feel oddly brittle.

While tuning retrieval for an enterprise-style assistant, that pattern showed up clearly. Missing evidence pushed the team to raise retrieval depth. Offline retrieval scores rose. Recall climbed. Paper metrics said “better.” Generated answers got worse.

Extra context was not helping. Sometimes it hurt. Relevant passages arrived tangled with duplicates, weak neighbors, and occasional conflicts.

That experience reframed evaluation. The question stopped being “did we pull the right file?” and became “can each stage of the pipeline show it is doing its job?”

Debugging live RAG systems quickly separates several axes that demos collapse into one score: how well you retrieve, how well you rank, how good the answer is, whether claims stay grounded, how long users wait, and what each request costs. The follow-up everyone asks — how should we pick chunk size and top-k? — no longer deserves a canned pair of numbers.

Treat the problem as offline optimization: invent ground truth, measure the right signals, name the trade-offs, check unseen queries, then confirm the win still holds in production.

Six measures are especially useful: Recall@K, nDCG, MRR, Faithfulness, Latency, and Cost. Each exposes a different failure. Together they explain how a prototype becomes a system you can operate.

Start with retrieval

When answers degrade, go back to the front of the pipeline before rewriting prompts, swapping models, or bolting on agents. Ask the blunt question: is the evidence even being found? A model cannot use material that never entered the context window.

1. Recall@K — coverage of needed evidence

Recall@K asks what fraction of truly relevant items appear among the top K hits.

Take an ITSM bot facing: “Payment service failed after a database failover — which recovery steps should I run?” Suppose five facts are required. The retriever’s top five include only four of them. Recall@5 is 0.80.

That single number already steers debugging. The generator might not be the villain; twenty percent of the needed evidence never arrived. Rule of thumb: improve what the model sees before blaming how it writes.

It also explains why a fixed top_k = 5 is not sacred. Moving K from 5 to 10 and watching Recall climb from 0.80 to 0.94 suggests the index contains the material but the cut is too shallow. Raising K also risks flooding the prompt with noise — which is why ranking metrics come next.

2. nDCG — are the best hits near the top?

Coverage alone is incomplete. Position matters.

Two systems can reclaim the same runbook. One places it first, then another useful procedure, then weaker chunks. The other buries the runbook at rank eight under loosely related pages. Similar recall; very different products.

nDCG (normalized Discounted Cumulative Gain) scores graded relevance and pays more for putting strong items early. The classic discounted-gain work by Järvelin and Kekäläinen exists for exactly that reason.

Reranking makes this concrete for RAG. Retrieve twenty candidates, keep five for the model, and the order of the twenty decides whether the five are any good. High recall with weak ranking can leave the right passage in the pool yet outside the final prompt.

Recall checks discovery. nDCG checks intelligent ordering.

3. MRR — how soon is the first useful hit?

Mean Reciprocal Rank zooms in on the first relevant result:

[
RR = \frac{1}{\text{rank of first relevant result}}
]

Rank 1 → RR 1.0. Rank 5 → RR 0.2. Averaged over a query set, MRR is a sharp signal when the first good hit dominates the experience.

Interactive assistants often stop after the first strong passage. A retriever that hides the right runbook at rank nine can look healthy on Recall@20 and still feel broken in the UI.

When “more context” backfires

After retrieval numbers improved, answer quality still fell. The model was swimming in redundant and conflicting text. Ranking metrics explain the trap: larger K without better order dumps noise into the prompt. That leads to grounding checks.

4. Faithfulness — claims tied to retrieved text

Faithfulness asks whether the answer’s claims are supported by the retrieved context. Smooth prose that invents a step, invents a threshold, or merges two policies into a third fails faithfulness even when it sounds confident.

Keep faithfulness separate from correctness.

Suppose the corpus still contains an obsolete password rule — expire every 60 days. The model retrieves it and repeats it. The reply can be fully faithful to the retrieved page and still wrong against the intended source of truth.

  • Faithfulness: supported by what was retrieved?
  • Correctness: true against the intended policy?

That split matters whenever knowledge bases drift under a live system.

5. Latency — quality users can wait for

Offline quality is useless if the product cannot afford the wait.

Compare two setups. A hits 91% answer accuracy with 120 ms retrieval latency. B hits 93% at 650 ms. B wins on accuracy alone. Product constraints decide whether B is actually better. Interactive assistants with hard response budgets may reject that delay, and retrieval is only one slice of total time.

A live request can include rewrite, retrieve, rerank, assemble context, and generate. Measure stages and the end-to-end path. Prefer distributions over averages. A 400 ms mean with a 1.8 s P95 feels nothing like a system whose P50/P95/P99 stay low.

Put latency into the evaluation plan from day one — not as an ops metric after the architecture freezes.

6. Cost — what happens at millions of queries

Economics show up the moment a prototype becomes a product.

Raising top-k touches more than recall. It can grow reranker work, final context size, input tokens, latency, and infrastructure burn.

If chunks average 600 tokens, then:

Final K = 5

injects about 3,000 retrieved tokens. Jump to:

Final K = 15

and you are nearer 9,000 — triple the retrieved mass. Tiny traffic hides the bill. Millions of requests turn it into a product choice. Top-k is simultaneously a quality, latency, and cost knob.

What “tune chunk size and top-k” really means

The interview question is not asking for two magic numbers. It is asking for an experimental design.

Build roughly 200 representative queries across troubleshooting, procedures, incidents/RCA, policy, ambiguity, multi-hop, and edge cases. Let domain experts judge relevance.

Sweep chunk sizes like:

Chunk sizes:
256
512
1024
2048

and overlap / candidate-K grids like:

Overlap:
64
128
256Candidate K:
5
10
20

A 4 × 3 × 3 grid yields about 36 configs. Score each with more than one number:

Recall@K
nDCG@K
MRR
Answer correctness
Faithfulness
Latency
Cost

Then pick the quality / latency / cost frontier — not the max recall or F1 by default.

Counterintuitive experiment outcomes

Suppose three configs land like this:

  • A: Recall@10 0.94, nDCG@10 0.81, accuracy 0.89, faithfulness 0.95, P95 510 ms, $0.025
  • B: Recall@10 0.91, nDCG@10 0.88, accuracy 0.94, faithfulness 0.96, P95 560 ms, $0.027
  • C: Recall@10 0.97, nDCG@10 0.79, accuracy 0.86, faithfulness 0.76, P95 820 ms, $0.034

Recall alone crowns C. C produces the weakest answers — extra retrieval apparently harms generation. B lacks top recall yet leads on nDCG, accuracy, and faithfulness with only a mild latency/cost penalty. That is the config worth digging into, and why “highest retrieval score wins” is a bad universal rule. Optimize for the product’s real objective.

Failure → next investigation

  • Weak Recall@K → chunking, embeddings, index, metadata filters, query shaping, retrieval strategy
  • Strong recall, weak nDCG → ranking / reranking
  • Strong retrieval, weak answer accuracy → context assembly, ordering, prompts, generation
  • Plausible but unsupported answers → faithfulness / grounding
  • Good quality, bad latency → profile every stage
  • Good quality, bad cost → candidate K, final K, chunk size, compression, caching, model choice, token efficiency

That checklist turns debugging into a procedure.

Keep the loop alive in production

A one-shot benchmark is not the end state. You want continuous evaluation. Live traffic diverges from curated sets: phrasing drifts, documents change, policies update, services appear, edge cases arrive. Grow the evaluation set from production failures.

Mature scorecard

No lone metric tells the whole story. A useful card tracks coverage, ranking, correctness, faithfulness, latency percentiles, and unit economics together.

Ask before you tune

When someone demands a chunk size and a top-k, answer with questions first: what are we optimizing, what does the labeled set look like, and what does a retrieval miss cost? With those answers, configs become experimental results.

One outcome might be:

512 tokens
128 overlap
Candidate K = 20
Final K = 5

Another:

1024 tokens
64 overlap
Candidate K = 10
Final K = 4

Semantic chunking might beat both. A reranker might make final K matter more than candidate K.

Do not trust a universal value without measurement. A RAG system is not “good” because it retrieves more, retrieves faster, or sounds convincing. It is good when it finds the right evidence, ranks it well, feeds the model the right amount of context, produces answers that are both correct and grounded, and stays inside the product’s latency and cost envelope.

That gap — between wiring a pipeline and engineering a production system — is the whole point.

Stop guessing the knobs

No mythical chunk length, no one-size top-k, and no lone metric can declare a deployment ready. Better Recall@K can still hurt answers. Strong retrieval can still permit unsupported claims. High accuracy can still fail if latency or cost explode at scale.

Balance coverage, ranking, grounded generation, latency, and cost for the workload you actually serve.

Prefer: ground truth → retrieval benchmarks → ranking checks → end-to-end answer quality → latency/cost validation → unseen holds → production monitoring → feed failures back into the set.

Then chunk_size=512 or top_k=5 is evidence, not folklore.

Putting the six metrics on one page

A workable scorecard for a weekly RAG review might look like this:

  • Recall@K and MRR for “did we find evidence, and how soon?”
  • nDCG for “did we put the best evidence first?”
  • Faithfulness and answer correctness for “did generation stay honest and true?”
  • P50/P95 latency for “can users wait?”
  • Cost per successful answer for “can finance wait?”

Review them together. A spike in Recall@K with a drop in faithfulness is not a win. A latency cut that collapses nDCG is not a win. The point of the six-metric frame is to make those trade-offs visible instead of burying them inside a single F1 number.

Ground truth is the scarce resource

Metrics are only as good as the labels underneath them. If domain experts never judged which passages were relevant, Recall@K is theater. If nobody marked graded relevance, nDCG collapses into a noisy binary. If faithfulness judges are inconsistent, grounding scores drift.

Budget time for labeling the same way you budget time for embedding experiments. Two hundred carefully judged queries usually teach more than two thousand unlabeled ones. Refresh that set from production tickets: every “wrong deductible,” “outdated password rule,” and “missed runbook” is a candidate case.

From experiment to change control

When a configuration wins offline, promote it like any other production change. Record the grid you swept, the winning settings, the hold-out results, and the latency/cost envelope you accepted. Then watch the same metrics on live traffic for a soak period. If production queries diverge, feed the failures back into the labeled set and rerun the grid. That loop — not a blog-post default for chunk size — is what separates a tuned system from a lucky demo.

Why prototypes lie

Demo notebooks usually fix the query set, freeze the corpus, and hide latency behind a single call. Production does the opposite: queries drift, documents churn, and users abandon slow answers. That is why a configuration that looked excellent on a static benchmark can feel unreliable after launch. The six metrics above are a way to make those hidden dimensions explicit before you ship, and to keep them explicit after you ship.

When someone asks for a default chunk size and a default top-k, translate the request into an experiment plan. Name the objective. Name the labeled set. Name the latency budget. Name the cost ceiling. Then let the grid produce the numbers. Anything else is guesswork dressed as engineering.

Keep the scorecard visible in the weekly review so trade-offs stay explicit for the whole team.