FIFTY.DEV
All field notesJournal /The retriever got the wrong question
AI · RAG2026.08.21 · 11 min

The retriever got the wrong question.

Clarify, rewrite, split, or chain before retrieval.

RAG without the magicPart 6 of 8

RAG, without the magic · Part 6 of 8.

01 / DIAGNOSISThe retriever got the wrong question

four illustrative requests expose four different shapes before any treatment is chosen:

  • leave could mean entitlement, process, sick leave, or parental leave.
  • PTO carryover may name a concept the source calls annual-leave rollover.
  • compare parental leave and sick leave contains two complete information needs that can be investigated independently.
  • which manager approved the project whose budget changed? contains a dependency: one sourced result is needed before the next question can be formed.

the same preprocessing step would not preserve intent across all four.

a user types one word into a retrieval-augmented generation system:

leave

the corpus contains an annual-leave policy, a sick-leave policy, a parental-leave policy, and a page explaining how to request time off. the retriever can search all of them. that is not the same as knowing which one the user meant.

the useful first move is not expansion. it is one question:

do you mean your leave entitlement, how to request time off, or a specific kind of leave?

those interpretations lead to different evidence and different answers. choosing one silently would make the system sound precise while changing the task underneath the user.

the wording is not defective. the system lacks enough intent.

this is the query-processing problem. a strong retriever can still miss when the search task does not represent the user's information need. before retrieval, the system may need to leave the question alone, clarify it, rewrite its wording, split independent questions, or retrieve one fact before asking the next.

the rule for this part is small: choose the least invasive treatment that preserves intent. direct retrieval stays the default. every transformation is a hypothesis, not an upgrade.

FIG. 01 — clarification before transformation when the original query lacks enough intentAMBIGUITY / CLARIFY

02 / TRANSFORMRewrite the wording, or rewrite the task

query processing becomes easier to reason about when it is split into two jobs.

rewrite the wording when there is still one information need. the user and the corpus may use different vocabulary, or several equivalent formulations may reach different candidate passages. query expansion, multi-query retrieval, and Hypothetical Document Embeddings, or HyDE, live on this side.

rewrite the task when the request contains separate information needs or a dependency. decomposition creates independent searches. multi-hop retrieval creates a sequence in which a later search depends on sourced evidence from an earlier one.

two exits sit before both jobs.

direct retrieval is correct when the question is specific, aligned with corpus vocabulary, and answerable in one search.

clarification is correct when plausible meanings would lead to materially different sources or answers.

that distinction matters because a clean rewrite can still be wrong.

FIG. 02 — rewrite the wording, rewrite the task, or leave the query aloneLEAST INVASIVE TREATMENT

03 / EXPANDBridge a clear vocabulary mismatch

suppose the intent is already clear and the user writes:

PTO

the policy corpus says paid time off, vacation days, and annual leave.

this is a vocabulary mismatch. query expansion adds reviewed terms that express the same intended concept in corpus language. a maintained domain vocabulary is an inspectable starting point for stable acronyms and aliases. it is attributable and easy to remove when an entry goes stale, but its safety still depends on freshness, ambiguity, locale, and scope.

a model can propose additions when the vocabulary is less predictable. those additions are derived data. log which prompt, model, rule, or dictionary produced each term, then compare the proposal with the original information need. an added term can improve candidate recall on a tested slice. it can also pull retrieval toward a related but different policy.[2][3]

keep the original query beside the expansion:

original: PTO
reviewed corpus terms: paid time off | vacation days | annual leave

do not replace the original field with the rewritten string. the original is the immutable intent anchor.

now return to leave.

a rewrite turns it into:

annual leave entitlement

retrieval may look cleaner. the system has also guessed twice. it chose annual leave instead of sick leave or parental leave, then chose entitlement instead of the request process. a plausible guess has become fake precision.

this is the failed rewrite seam. expansion is useful after intent is clear. it is not a substitute for clarification.

04 / MULTI-QUERYOne intent, several complete formulations

query expansion changes the terms inside one search representation. multi-query retrieval creates alternate complete formulations of the same information need.

start with:

How do I request PTO?

keep that original visible. then create variants with explicit purposes, for example:

policy wording: PTO request process and procedure
application wording: How to apply for paid time off
form wording: Vacation day request form submission

each variant must preserve the request action. a variant about PTO entitlement would not be another angle. it would be a different question.

issue the eligible searches under the same authorization scope. independent searches can run concurrently. that means wall time does not have to grow as the arithmetic sum of every search. the work still grows: there are more requests, more search compute, more candidates, more rate-limit pressure, and more fusion work. concurrency changes elapsed time, not the amount of work.

the same passage may arrive through the original and several variants. deduplicate before reranking and context assembly by a stable corpus identity, not by loose text similarity. a useful identity can include the document, document version, chunk, parent, tenant or security partition, and corpus version. two policy versions are not duplicates merely because their text overlaps.

deduplication collapses one identity. it does not erase how that identity was found. the surviving candidate retains every query path, rank, score, source version, and authorization decision that led to it.

after that, merge the ranked candidate lists using the fusion method established in Part 5. Part 5 owns the ranking mechanics. the point here is the flow:

original → variants → concurrent eligible searches → stable-identity deduplication → Part 5 fusion → authorized evidence

several variants can repeat the same mistaken assumption. paraphrase diversity is not independent evidence.

FIG. 03 — one intent, several search paths, with identity and provenance preservedMULTI-QUERY

05 / HYDEAn answer-shaped probe that is not an answer

sometimes the mismatch is structural. the user asks a short question. the corpus stores detailed, declarative passages.

HyDE generates a hypothetical document, embeds that synthetic text, and uses the embedding to retrieve nearby real corpus documents.[5] think of it as an answer-shaped search probe.

the label needs to stay attached:

SYNTHETIC RETRIEVAL PROBE — NOT EVIDENCE

imagine the probe says:

employees receive 25 vacation days and submit requests through the leave portal.

the 25 vacation days detail may be invented. the probe can help locate a corpus neighborhood. it cannot support the answer, appear in a citation, authorize a decision, or become a fact in a later hop. only retrieved, authorized source passages with provenance can do that.

the boundary is operational, not rhetorical. mark the probe as synthetic in traces and stored artifacts. keep it out of the evidence set and answer context. record its protected text or hash, generation details, embedding version, corpus version, retrieved candidate identities, and evidence=false.

if the probe finds no trustworthy source, stop or use the product's declared fallback. do not answer from the probe.

a detailed query that already matches the corpus may not need HyDE. again, the technique follows an observed failure. it does not follow the mere presence of a question mark.

FIG. 04 — a synthetic search probe stops at the evidence boundaryHYDE / NOT EVIDENCE

06 / PLANIndependent branches are not dependent hops

consider this request:

Compare our vacation policy with our sick-leave policy.

it contains two independently answerable questions:

what does the current vacation policy say?
what does the current sick-leave policy say?

this is query decomposition. each branch can retrieve its best sources without requiring one document to mention both policies.[6] the branches may run concurrently because neither answer is needed to form the other search.

they are still parts of one request. each subquery keeps the original employer or tenant, employee class, jurisdiction, effective date, comparison criteria, and authorization scope. authorize each branch separately. cite each side separately. during synthesis, compare only dimensions supported by evidence on both sides. if one policy does not state a dimension, name the gap instead of filling it.

decomposition is a reliability tool when compound wording causes candidate misses or uneven evidence. it is not mandatory for every comparison. a direct retriever may already return the right separate sources.

now consider:

What is the budget for the project John leads?

the second search cannot be written safely yet. the project name is missing.

the first hop asks which project John leads. a retrieved, authorized source must identify the person, project, and relevant time. only that supported project identity can become a search key for the budget hop. the second hop then retrieves the approved budget for that project.[7]

that dependency makes it multi-hop retrieval.

if the first hop finds several people named John, several current projects, conflicting assignments, or no authorized evidence, stop. clarify or abstain. do not let a model guess choose the next query.

provenance crosses the hop boundary. record the query, the exact prior source fact that justified the next query, the source and chunk version, the authorization decision, the accepted or unresolved state, and the stop or continue reason. the final answer cites both the source that links John to the project and the source that states the budget.

discovering a project identifier does not grant access to its finance records.

FIG. 05 — independent branches can run together; dependent hops cannotDECOMPOSE / HOP

07 / ROUTEA compact query-level router

the router for one incoming question can stay compact:

  • direct: the request is specific, corpus-aligned, and has one information need. search it unchanged. stop when cited evidence satisfies the request, or when evidence is insufficient.
  • clarify: plausible meanings lead to materially different evidence or answers. ask one focused question and wait. stop until ambiguity is resolved.
  • rewrite or expand: intent is clear but corpus vocabulary differs. retain the original, add reviewed equivalents, and reject terms that change the subject, relation, constraints, or answer form. stop or roll back when drift appears or retrieval does not improve on the tested slice.
  • multi-query: one clear intent benefits from alternate complete formulations. run eligible variants, deduplicate by stable identity, then use Part 5 fusion. stop when added variants produce no useful new evidence or the configured budget is exhausted.
  • decompose: the request contains independent information needs. preserve shared constraints, authorize every branch, and synthesize from separately cited evidence. stop when a branch lacks enough evidence rather than inventing symmetry.
  • multi-hop: a later query depends on an earlier sourced fact. carry citation, confidence, and authorization through each dependency. stop on ambiguity, conflict, repeated state, insufficient evidence, policy denial, or exhausted budget.

if the router is uncertain, preserve the original and clarify, abstain, or follow the product's declared fallback. choosing a route is a policy or classification decision. it is not proof that the route is correct.

this is query-level routing. it chooses how one request reaches retrieval. it does not define the durable architecture behind every path. that belongs to Part 7.

FIG. 06 — route the query, not the architectureQUERY-LEVEL ROUTER

08 / GUARDRAILSThe guardrails travel with the query

query processing changes what the system searches for. it must never change what the caller is allowed to search or read.

every rewrite, variant, decomposition branch, HyDE search, and hop inherits the original requester's trusted authorization envelope. apply it on every transformed request. revalidate each selected object after retrieval and again before it enters model context or leaves the system. unknown identity, tenant, object permission, or policy state fails closed.[9]

keep one linked trace:

  • the immutable original query or a protected reference to it
  • the clarification question and answer, when used
  • transformation type, parent transformation, purpose, and transformed text
  • prompt, model, rule, or dictionary version
  • authorization-scope reference and decision for every branch or hop
  • corpus, index, retriever, embedding, document, and chunk versions
  • stable candidate identities, per-path ranks or scores, and the deduplication map
  • selected evidence and citation provenance
  • confidence or sufficiency decision
  • stop, abstain, or escalation reason

query text can contain personal, confidential, or regulated data. retaining the original does not mean copying raw text into every log. the logging policy can store a protected reference, minimize fields, redact or tokenize sensitive values, separate operational metadata from content, encrypt traces, restrict access, and define retention and deletion. denied source text does not belong in observability systems.[6]

stop when the information need is satisfied, evidence is insufficient or conflicting, ambiguity remains material, authorization blocks the next action, confidence falls below the product's declared threshold, a state repeats, or the configured time or compute budget is exhausted. there is no universal query count or hop count.

start with the unchanged question. add one treatment for one observed failure. keep it only when measured retrieval improves without degrading intent, authorization, or provenance. Part 8 owns how that claim is tested.

one unresolved question remains. a router can choose a path only when the system has the right bounded paths and controls available. Part 7 is about that architecture.

reference notes

[1] Asking Clarifying Questions in Open-Domain Information-Seeking Conversations: https://arxiv.org/html/1907.06554

[2] Introduction to Information Retrieval, Query expansion: https://nlp.stanford.edu/IR-book/html/htmledition/query-expansion-1.html

[3] Query-drift prevention for robust query expansion: https://dl.acm.org/doi/10.1145/1390334.1390524

[4] RAG-Fusion: a New Take on Retrieval-Augmented Generation: https://arxiv.org/html/2402.03367

[5] Precise Zero-Shot Dense Retrieval without Relevance Labels: https://arxiv.org/html/2212.10496

[6] Microsoft Azure Architecture Center, Information retrieval: https://learn.microsoft.com/en-us/azure/architecture/ai-ml/guide/rag/rag-information-retrieval

[7] Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions: https://arxiv.org/html/2212.10509

[8] Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity: https://arxiv.org/html/2403.14403

[9] OWASP Authorization Cheat Sheet: https://cheatsheetseries.owasp.org/cheatsheets/Authorization_Cheat_Sheet.html

RAG without the magic

Continue reading

View all 8 parts →
FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 6 · THE QUESTION●2026.08.21●FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 6 · THE QUESTION●2026.08.21●