FIFTY.DEV
All field notesJournal /The answer failed. Which half broke?
AI · RAG2026.08.21 · 17 min

The answer failed. Which half broke?

Evaluate both stages. Keep the trace.

RAG without the magicPart 8 of 8

RAG, without the magic · Part 8 of 8.

a user asks the HR assistant whether part-time employees qualify for parental leave.

the answer is wrong.

that tells us almost nothing.

maybe retrieval never found the policy. maybe it found the policy, ranked it below weaker documents, then dropped it when the context was assembled. or maybe the correct policy reached the model and the model ignored it, mixed it with another rule, or answered when it should have abstained.

one bad answer. several different repairs.

a RAG system can be inspected like a two-stage assembly line. the warehouse retrieves candidate parts. the assembly floor builds an answer from the parts it receives. the picture is diagnostic rather than a literal universal architecture: context assembly sits at the seam, and the terminal result may be an answer, abstention, or escalation. the stages can fail independently, so every investigation starts with two questions:

  1. did retrieval find it?
  2. did the model use it?

do not begin with a dashboard. begin with a record that lets you replay the failure.

FIG. 01 — two stages, four outcomes: diagnose ownership before repairFAILURE OWNERSHIP

01 / RECORDDefine the eval record first

an evaluation metric without an inspectable row cannot explain which stage failed.

a useful eval record keeps the query, the evidence that should have been found, the candidates that were retrieved, the exact context supplied to the model, the answer it produced, the labels applied, and the versions that produced all of it.

for the opening question, an illustrative replayable row might say:

  • expected evidence: the current parental-leave policy and its part-time eligibility clause.
  • retrieved candidates: an older full-time policy, a benefits summary, and the current policy at a rank below the context cutoff.
  • supplied context: only the first two candidates after filtering and truncation.
  • observed answer: a confident denial with no support for the part-time rule.
  • diagnosis: retrieval found the required source but context assembly omitted it; generation then answered without sufficient evidence instead of abstaining.

that trace earns a specific repair: inspect ranking and cutoff behavior first, then keep abstention as a generation guardrail. it does not justify changing every stage at once.

at minimum, record:

record id and timestamp
query text, conversation state, source, and task, risk, answerability, and relevant slices
caller role and tenant, pseudonymized where applicable, plus authorization-policy version
evaluated unit: chunk, passage, document, or source
accepted relevance labels and graded judgments
label rubric and version, effective date, labeler role, evidence inspected, confidence, disagreement, and adjudication
expected evidence ids or passages, their versions, required source coverage, and required answer points with severity
acceptable abstention or escalation state
ranked retrieved candidates with scores, versions, eligibility decisions, and retrieval-path or query-variant provenance
exact supplied context after filtering, reranking, expansion, compression, and truncation
observed answer, citations, and claim-to-citation links
retrieval, generation, citation, abstention, security, and other failure labels
primary metric and guardrails
metric implementation or library version or commit wherever metrics are reported
eval-set, corpus and document, parser and chunker, embedding, retriever and index, filter, reranker, and context-builder versions
prompt, generator or model, judge, and decoding versions
code and configuration, experiment, and dated price-sheet versions
privacy, authorization, and security flags
reviewer or judge identity and configuration
outcome

the evaluated unit matters. a relevant chunk is not the same thing as a relevant source. recall@5 over chunks and recall@5 over sources answer different questions even when they use the same query.

the versions matter too. if the prompt changed, the corpus moved, the index was rebuilt, or the judge model was replaced, the score belongs to a different measuring setup. preserve enough of that setup to run the failure again.

now the two questions can be answered with evidence.

FIG. 02 — keep the failure replayable with a versioned eval recordVERSIONED EVAL RECORD

02 / RETRIEVALDid retrieval find it?

retrieval metrics are not a menu to collect. each one belongs to a failure question.

before using any metric with @K, define two things:

  • K: how many ranked results are being inspected
  • the unit: chunk, passage, document, source, or another declared item

choose K from the real retrieval stage and context budget. also record what happens with duplicates, ties, unjudged items, and queries that return fewer than K results. the implementation is part of the result, not a footnote someone can guess later.

Did retrieval miss necessary evidence? use Recall@K

suppose the eval says three source-level policies are relevant to a query. the retriever returns ten sources, and two of the three relevant sources appear in those ten.

Recall@10 over sources
= relevant labeled sources in the first 10
  / all labeled relevant sources for the query
= 2 / 3

Recall@K asks how much of the accepted relevant set appeared within the cutoff. it requires relevance labels. if the accepted set is incomplete, the score is conditional on those judgments. it does not prove that every relevant source in the world was found.

use it when missing evidence is the failure you are trying to fix.

Did retrieval add distracting material? use Precision@K

now take the same first ten source-level results. two are labeled relevant:

Precision@10 over sources
= relevant labeled sources in the first 10 / 10
= 2 / 10

that is the usual fixed-cutoff convention. if your implementation divides by the number actually returned when fewer than K appear, say so.

Precision@K asks how concentrated the top K is with relevant units. it also requires relevance labels. it does not tell you whether every necessary source was found.

use it when noisy context may be distracting the generator.

Was the first usable result buried? use MRR

for one query, reciprocal rank is:

1 / rank of the first relevant result

if no relevant result appears in the evaluated list, the reciprocal rank is zero under the usual convention. Mean Reciprocal Rank, or MRR, is the arithmetic mean of that value across the evaluated queries.

MRR answers one narrow question: how early did the first relevant result appear?

it ignores every relevant result after the first. it does not measure whole-list quality, completeness, or source coverage.

this catches a quiet failure. the right policy may be retrieved at rank 18, then disappear when only the first 10 source-level results are passed onward. retrieval technically found it. the generator never saw it.

Are the most useful results near the top? use NDCG@K

sometimes relevance is graded rather than binary. one source may answer the question directly, another may supply a necessary exception, and a third may be related but weak.

Normalized Discounted Cumulative Gain, or NDCG@K, evaluates a ranked list using declared relevance grades. more useful units receive more gain. lower ranks receive a discount. the observed ranking is normalized against the ideal ranking for the same judged items.

declare the cutoff, evaluated unit, label scale, gain mapping, discount, and behavior when the ideal gain is zero. a scale from 0 to 3 can be a useful example, where 0 means not relevant and 3 means highly relevant. it is an example, not a standard. different scales and gains produce different measurements.

NDCG@K asks whether more useful units appear earlier than less useful ones. unlike MRR, it can reflect the graded quality of more than the first relevant result.

the map is simple:

  • missing necessary evidence: Recall@K
  • too much noise in the cutoff: Precision@K
  • first relevant result appears too late: MRR
  • graded ranking quality across the cutoff: NDCG@K

the metric follows the failure question.

FIG. 03 — recall finds misses, Precision exposes noise, reciprocal rank checks the first relevant result per row, MRR averages it across rows, and NDCG evaluates declared graded ranking qualityRETRIEVAL DIAGNOSTICS

03 / GENERATIONDid the model use it?

if the warehouse supplied the right parts, inspect the assembly floor.

Faithfulness means support from supplied context

faithfulness, or context support, asks whether the answer's checkable claims are supported by the evidence actually supplied to the generator.[2]

it does not prove that a claim is true in the world. it does not prove the source is current, authorized, complete, or authoritative. it also cannot prove where the model first learned a phrase or fact. it checks support against a declared context and judging procedure. nothing more mystical is happening.

Answer relevance means addressing the request

answer relevance asks whether the response addresses the user's question. it is not correctness.

a direct answer can be wrong. a well-supported answer can avoid the requested task.

consider this request:

what steps do i follow to submit a purchase request, and who approves it?

the supplied context contains the process. the model replies:

the company maintains a purchase-request policy to support consistent procurement and responsible spending.

every sentence may be supported by the context. the answer is still useless. it reads like a PR statement. it gives no steps and names no approver.

that is faithful but irrelevant and incomplete.

Completeness means covering required points

completeness asks whether the answer includes the accepted required points and necessary caveats. for the purchase request, the eval row might require:

  • the submission steps
  • the approval role
  • the exception for requests above a declared policy limit
  • an abstention or escalation if the applicable policy version cannot be verified

length is not completeness. a short answer can cover every required point. a long answer can miss the one that matters.

Citation accuracy is not citation presence

citation coverage asks whether claims that need citations have them. citation accuracy asks whether each cited source supports the claim attached to it. source coverage asks whether the answer uses enough of the required sources for a multi-source task.[3]

keep these separate. a citation can exist and still point to the wrong passage, an obsolete policy, an unauthorized source, or evidence that does not support the claim.

the diagnostic stays the same:

  • if retrieval did not supply the evidence, fix retrieval, indexing, ranking, filtering, or context assembly
  • if retrieval supplied the evidence and the answer failed, fix generation, prompting, context use, citation behavior, or abstention

polishing the answer layer will not recover a policy that never reached it.

FIG. 04 — supported can still be irrelevant, incomplete, or poorly citedGENERATION DIAGNOSTICS

04 / LABELSThe exam needs an answer key

evaluation has a cold-start problem. you cannot grade an exam without an answer key.

in this setting, “ground truth” means the accepted labels and judgments used for the eval. it is not metaphysical truth. record the rubric, label version, evidence inspected, labeler role, confidence, disagreement, adjudication, and effective date.

build the test set from three sources. keep the source tag on every row.

Expert-authored and expert-labeled cases

experts are useful for high-value, ambiguous, policy-sensitive, and high-risk cases. they can identify required evidence and costly omissions that a generic generator may miss.

their labels are not automatically unanimous. expert sets are slower, smaller, and shaped by institutional assumptions. use a written rubric, double-label a risk-based subset, adjudicate load-bearing disagreements, and retain legitimate dissent.

Synthetic cases

synthetic queries create controlled breadth. they are useful for rare edge cases, permission boundaries, known retrieval misses, unanswerable questions, prompt-injection probes, and poisoned-document tests.

they inherit the generator, prompt, seed document, and template. they may copy source wording, overproduce clean answerable questions, and miss the ambiguity of real requests. pin the generation setup, tag the rows, audit a sample, and do not use synthetic cases as the only acceptance set.[4]

Production-derived cases

production queries, support escalations, incidents, corrections, and sampled conversations bring real wording and observed failures.

they also bring selection bias. logs reflect current users, current permissions, current product flows, logging choices, and the people who chose to complain. they miss silent dissatisfaction, abandoned sessions, blocked requests, non-users, and future traffic. production logs are not automatically representative.

apply privacy review before a row enters the eval set. minimize and redact content, define consent or another documented basis where applicable, restrict access, set retention and deletion rules, and deduplicate near-copies before splitting development and held-out sets. do not tune and report on the same cases.[1]

the useful mix is not one source winning. it is expert validity, synthetic control, and production language, separately tagged so their biases stay visible.

05 / JUDGEAn LLM judge is an instrument

human review does not scale to every candidate answer. an LLM-as-judge can help triage context support, answer relevance, completeness, and citation support. treat it like a measuring instrument, not an oracle.

before relying on it:

  1. write a rubric with positive, negative, and borderline examples
  2. calibrate against an accepted human-labeled sample that includes ordinary, refusal, long-context, multilingual, citation, and adversarial cases
  3. pin and log the judge model, prompt, examples, decoding settings, evidence order, output schema, parser, retries, and version
  4. check reproducibility with repeated runs where nondeterminism remains
  5. probe known biases, including answer order, response length, confident style, familiar wording, and preference for the system being evaluated
  6. inspect confusion and severe false passes by slice, not only an average score
  7. route disagreements, ambiguous cases, severe failures, and a random audit sample to humans

a judge change is a measurement change. freeze the judge for an experiment. recalibrate when its model, prompt, parser, task mix, language mix, or observed agreement changes.[5]

keep deterministic checks outside the judge where possible: source ids, citation existence, authorization decisions, empty retrieval, timeouts, index versions, and schema validity do not need a model's opinion.

FIG. 05 — combine expert, synthetic, and production-derived cases while preserving their biasesEVAL-SET MIX

06 / OFFLINEOffline comparison is not an A/B test

an offline benchmark or ablation runs a baseline and a candidate on a fixed, versioned eval set. the same rows make failures inspectable and paired differences visible.

before the run, declare:

  • the failure question and evaluated population
  • exactly what differs between baseline and candidate
  • the primary metric, or a small ordered family, that decides the intended win
  • the minimum practically important effect and decision rule
  • guardrails that must not regress
  • relevant slices
  • sample-size assumptions
  • how uncertainty will be reported, including whether an observed difference is large enough to distinguish from expected sampling noise under the declared test
  • exclusions, missing data, repeated checks before the planned end, and corrections when several decisions are tested at once

changing one factor is a useful debugging default, not a law. a deployable bundle can be tested as a bundle. state what changed.

there is no universal minimum sample size. it depends on the baseline rate or variance; the smallest effect worth detecting; the tolerated chance of a false alarm; target power, meaning the planned chance of detecting that effect if it is real; allocation, meaning how observations are split between variants; clustering or repeated measures, where observations from the same user or group are not independent; exclusions; and the number of decisions being made.[6]

report effect size and uncertainty, not only a “significant” label. preserve query-level results and slices. an aggregate mean can hide a failure for one language, permission scope, source family, or high-risk task.

an online A/B test is different. it assigns eligible live traffic, users, sessions, requests, tenants, or another declared exposure unit to variants. it measures user or product outcomes under real operation. assignment, eligibility, allocation, exposure logging, privacy, sample size, analysis, stop conditions, incident ownership, and rollback must be declared before exposure.

offline asks: did the candidate improve labeled behavior on this fixed set?

online asks: did assignment to the candidate change observed production outcomes?

do the first before accepting the risk of the second.

07 / REGRESSIONA retrieval win can make the answer worse

suppose policy queries often miss one required source. the team broadens retrieval. on the fixed offline set, Recall@10 over sources improves.

that is the intended win.

but the broader candidate set also introduces stale and tangential policy material. more of it survives into the supplied context. the generator starts mixing an old exception into current answers. some responses remain supported by something in the packet but answer the wrong policy question. answer relevance falls, citation support becomes less reliable, and the stale-source slice regresses.

the retrieval metric improved. the system got worse.

a predeclared primary metric and guardrails catch this without collecting every metric available:

  • primary metric: Recall@10 over sources for the targeted policy-query set
  • guardrails: answer relevance, faithfulness-to-context, citation accuracy and source coverage, stale-source failures, authorization leakage, latency, and cost

the primary metric explains the intended repair. the guardrails show whether the repair broke the assembly floor or the operating envelope.

08 / PRODUCTIONProduction changes the evidence

a fixed eval set tells you what happened on a known collection of questions. production adds new wording, new documents, new permissions, new failure paths, and new incentives. monitor it in phases, led by the failures the system can plausibly make.

separate three classes of signal:

  1. directly measured operations
  2. automated quality estimates, including calibrated judge output
  3. human or user evidence

they are not interchangeable.

Operations

record end-to-end and stage latency. where traffic supports stable estimates, use percentiles: p50 is the midpoint observation, p95 is slower than 95 percent of observations, and p99 is slower than 99 percent. separate successful, failed, timeout, fallback, first-token, and complete-response paths.[8]

record per-query and aggregate cost with the components that created it: tokens, model calls, retrieval and reranking operations, retries, cache behavior, tools, judge calls, and the dated price configuration. cost without its configuration drifts quietly.

watch empty retrieval, zero eligible evidence after authorization, post-filter depletion, duplicate collapse, timeouts, retries, fallback routes, and terminal outcomes. track abstention and escalation with coverage: correct supported answers, harmful answers, correct abstentions, false refusals, completed escalations, and unresolved cases.[14]

Corpus and citations

monitor source event time, ingestion lag, backlog, failed or partial index jobs, document and chunk counts, active index snapshot, delete propagation, permission changes, and stale-answer incidents. a successful job message does not prove every derived store is current.

monitor citation coverage separately from sampled citation support, source authority, effective date, and broken references. a citation icon is not evidence that the claim is supported.

Quality, slices, and drift

user feedback, corrections, support tickets, and complaints are noisy sampling signals. they are not labels by themselves. sample cases for human review based on risk, incidents, uncertainty, and relevant slices. do not impose a universal calendar.

where a calibrated judge is used, version it and monitor judge-human agreement. a drifting judge can make a stable system look better or worse.

compare only slices tied to a failure hypothesis or safety requirement: query class, answerability, language, product area, source type, permission scope, freshness band, tenant or role where privacy allows, and architecture route. keep slice definitions versioned. suppress or combine cells that create privacy risk.

drift is a change in inputs, corpus, labels, system behavior, or outcomes. a moving dashboard line is a clue, not a diagnosis. compare against declared reference windows, inspect affected slices, and relabel a sample before declaring quality drift.

Privacy and security

log with a purpose. minimize raw queries, answers, and evidence. redact or tokenize identifiers before persistence, restrict access, encrypt where required, define retention and deletion, preserve consent or opt-out rules, and review judge or vendor data transfer.[9]

security is a release gate, not an average score.

test the same query under allowed and denied roles. test revoked access through the index, cache, context builder, citations, fallback routes, query transformations, graph paths, tools, and generated response. authorization must be enforced independently at every boundary and fail closed when identity, permission, or tool scope is uncertain.[11]

seed direct and indirect prompt injection in queries, documents, hidden html or pdf text, metadata, retrieved snippets, citations, and tool output. retrieved text is untrusted data. RAG can carry an instruction into the model; it does not neutralize it.[10]

seed poisoned documents too: forged authority, stale or conflicting versions, duplicate authoritative-looking text, manipulated metadata, and malicious corpus updates. verify admission, provenance, precedence, quarantine, reindexing, rollback, answer support, and incident detection.[12]

when authorized evidence is insufficient, conflicting, stale, unavailable, or denied, the system may need to abstain, ask for clarification, or escalate. test that behavior as part of the product path. a refusal without a safe next step can be correct and still leave the user stranded.

09 / RELEASEShip reversibly

before exposing a new model, prompt, embedding, index, corpus snapshot, retrieval configuration, architecture route, or judge, canary the version on a bounded surface chosen for the product's risk and traffic.

declare the stop or rollback condition before launch. tie it to the baseline, primary metric, guardrails, and harm being prevented. keep the previous known-good versions and the deployment path back to them. test that rollback before the canary needs it.[13]

if an incident happens, preserve the trace. record impact, detection, timeline, versions, evidence, mitigation, rollback, contributing causes, and owned follow-up. label whether retrieval failed, generation failed, or the system should have abstained. add the relevant privacy and security notes. assign the case to a useful slice.

then turn the confirmed failure into a regression row after privacy review, labeling, deduplication, and versioning.

the warehouse changes. the assembly floor changes. the answer key changes with them. the loop stays useful only while each failure can be replayed and owned.

the next query can become the next test case only after privacy review, labeling, deduplication, and versioning.

FIG. 06 — test offline, expose carefully, learn from failure, and keep rollback realOPERATING LOOP

Reference notes

  1. Stanford, Introduction to Information Retrieval, “Evaluation”: https://nlp.stanford.edu/IR-book/pdf/08eval.pdf
  2. Es et al., RAGAS: Automated Evaluation of Retrieval Augmented Generation: https://arxiv.org/abs/2309.15217
  3. Gao et al., Enabling Large Language Models to Generate Text with Citations: https://arxiv.org/abs/2305.14627
  4. Saad-Falcon et al., ARES: An Automated Evaluation Framework for RAG Systems: https://arxiv.org/abs/2311.09476
  5. Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena: https://arxiv.org/abs/2306.05685
  6. Kohavi et al., Controlled experiments on the web: survey and practical guide: https://doi.org/10.1007/s10618-008-0114-1
  7. NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile: https://doi.org/10.6028/NIST.AI.600-1
  8. Google SRE, “Monitoring Distributed Systems”: https://sre.google/sre-book/monitoring-distributed-systems
  9. NIST, Privacy Framework: https://www.nist.gov/privacy-framework/privacy-framework
  10. OWASP, “Prompt Injection”: https://github.com/OWASP/www-project-top-10-for-large-language-model-applications/blob/99f4395589bdbd120ae961f9cd179e79d7f9b27f/2_0_vulns/LLM01_PromptInjection.md
  11. OWASP, “Vector and Embedding Weaknesses”: https://github.com/OWASP/www-project-top-10-for-large-language-model-applications/blob/99f4395589bdbd120ae961f9cd179e79d7f9b27f/2_0_vulns/LLM08_VectorAndEmbeddingWeaknesses.md
  12. OWASP, “Data and Model Poisoning”: https://github.com/OWASP/www-project-top-10-for-large-language-model-applications/blob/99f4395589bdbd120ae961f9cd179e79d7f9b27f/2_0_vulns/LLM04_DataModelPoisoning.md
  13. Google SRE, “Release Engineering” and “Postmortem Culture”: https://sre.google/sre-book/release-engineering and https://sre.google/sre-book/postmortem-culture
  14. Geifman and El-Yaniv, Selective Classification for Deep Neural Networks: https://arxiv.org/abs/1705.08500
RAG without the magic

Continue reading

View all 8 parts →
FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 8 · EVALUATION●2026.08.21●FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 8 · EVALUATION●2026.08.21●