FIFTY.DEV
All field notesJournal /Similar is not relevant
AI · RAG2026.08.21 · 10 min

Similar is not relevant.

Related passages are candidates. The answer still has to earn its place.

RAG without the magicPart 5 of 8

RAG, without the magic · Part 5 of 8.

01 / FAILUREThe missing condition is fourth

a new employee asks:

“how many vacation days do new employees get?”

the retrieval system returns this hypothetical ranking:

  1. “annual leave entitlement is 21 days per calendar year for all full-time employees.”
  2. “leave requests must be submitted through the HR portal at least two weeks in advance.”
  3. “the company holiday schedule includes ten public holidays annually, separate from personal leave allocation.”
  4. “employees joining mid-year receive prorated annual leave based on their start date.”

all four passages concern leave. only the fourth addresses the condition hidden inside “new employees.” it still cannot answer the numerical question alone: a complete answer needs the base entitlement, the proration rule, and the employee's start date. the search has found the subject and missed the question.

nothing mysterious broke. the first comparison did the job it was given: order related candidates. relatedness is not the same as answering. similarity is not relevance.

fixing that is not one upgrade. it is a sequence of responsibilities:

constrain → retrieve → merge → rerank → expand

each stage answers a different failure. each adds work and another way to lose useful evidence. the point is not to install the whole sequence. the point is to find which responsibility failed.

02 / CONSTRAINConstrain what is eligible

before comparing passages, the system can narrow which records are eligible for this retrieval operation.

think about filtering products on Amazon. category, price band, and rating do not appear because the store judged every product again when you clicked a filter. they are labels already stored on each product record. the filter reads those labels and removes records that do not qualify.

retrieval metadata works the same way. preparation may attach labels such as department, document type, intended audience, effective date, language, or version to each searchable passage. a question about a current HR policy might constrain the operation to records labelled department: HR, document_type: policy, and version: current before matching begins.

those labels change where the system may search. they do not make the matching inside that set more intelligent.

they can also be wrong. a missing label, a stale version, or an over-specific rule can remove the answer before any retriever sees it. filter selectivity and index implementation can change both latency and candidate recall. metadata filtering is an eligibility decision with costs to measure.

one boundary matters more than the rest: retrieval filtering is not authorization.

metadata says what an item is. authorization decides whether this person may perform this action on this object now. that decision must be independent, validated for every request and selected object, and fail closed when identity or permission is unknown. a tenant, role, or access_level label can help carry a trusted policy decision into retrieval. the label alone is not proof of permission. Part 2 owns that boundary.

FIG. 01 — retrieval eligibility is not an authorization boundaryLABELS ≠ PERMISSION

03 / RETRIEVERetrieve through complementary signals

once the eligible set is defined, the system needs candidates. one path rarely exposes every useful signal.

ask:

“what is form HR-101?”

a lexical path can locate the literal identifier HR-101. a representation-based path may instead return fluent passages about applying for leave that never mention the form. exact identifiers, clause numbers, error codes, and product symbols are where textual evidence often matters.

now ask for an entitlement using wording the policy does not use. the representation-based path may recover a paraphrase that a lexical path ranks poorly. lexical retrieval can still use stemming, fields, weighting, and configured synonyms, so this is not a law dividing exact questions from conceptual ones. it is a reminder that the paths fail differently.

running both paths is hybrid retrieval. operationally, that means applying the same eligibility policy, recording a ranked list from each path, normalizing candidate identity, deduplicating repeated passages, and forming a union. an extra path can also add noise. if neither input list contains the answering passage, the merge cannot invent it.

04 / MERGEMerge rankings without pretending the scores agree

the two paths return scores produced by different systems. a lexical score and a similarity score do not become comparable because they occupy adjacent columns.

ranked-list fusion avoids that direct comparison. it uses each candidate’s position in each list. a candidate near the top receives more credit than one near the bottom. contributions from every list are added. a passage that ranks reasonably well in both lists can therefore move above one that leads a single list and is absent from the other.

reciprocal rank fusion, or RRF, is a simple baseline built on that idea:

RRF(d) = Σ 1 / (k + rank_r(d))

for each ranked list containing passage d, take the reciprocal of a tunable constant k plus the passage’s position, then add the contributions. the constant dampens how strongly the first few positions dominate. the original paper used k = 60 after a pilot. that is historical context, not a magic number.

RRF is useful when component scores do not share a scale because it combines positions instead of raw values. it still needs decisions about the rank constant, input-list depth, candidate identity, ties, and path weighting. a passage outside a truncated list contributes nothing from that path. RRF also discards score gaps: a narrow lead and a wide lead both become rank one.

that makes RRF a useful, low-tuning baseline, not a universal default. compare it with each input path. a weighted score combination offers more direct control, but only after score direction and scale are normalized and a weight is selected on relevant labelled questions. neither fusion rule can recover a passage absent from every input list.

FIG. 02 — merge positions, then judge each candidate against the full questionHYBRID / RERANK

05 / RERANKCast wide. Then order the shortlist

candidate retrieval and reranking are two jobs with different operating limits.

the first stage casts a wide net over the eligible corpus. a common bi-encoder setup produces query and passage representations separately, which lets passage representations be prepared and indexed ahead of time. that makes broad retrieval practical, but the comparison sees two separately produced representations. it can identify the subject without fully testing how this passage answers this question.

the second stage orders the shortlist. a cross-encoder reranker reads the question and one candidate passage together. that joint view lets it weigh details such as “new employees,” “how many,” and “joining mid-year” against each passage rather than treating leave-related language as sufficient.

return to the vacation trace. the retriever placed the mid-year passage fourth. the reranker sees that “joining mid-year” is how this policy describes a new employee and moves that passage to position one. it improved the order because the proration condition was already present.

that moved passage is still not a complete answer. the numerical result needs the base allowance from passage one, the proration condition from passage four, and the employee's start date. the retriever decides what the reranker is allowed to see; a reranker can reorder a candidate set, but it cannot recover necessary evidence the first stage missed.

joint comparison costs latency and compute for every query-passage pair. that is why reranking usually operates on a shortlist rather than the whole corpus. there is no durable candidate count to copy from another system. candidate depth is a measured trade-off: a broader set can expose more answer-bearing passages, while also increasing fusion and reranking work. candidate depth, rerank depth, final evidence count, and final context tokens are separate budgets.

a dedicated reranker produces model scores used for ordering. the score is not automatically a probability, confidence, or calibrated measurement. a dedicated model can still capture nuanced query-passage relationships even when it emits no prose explanation.

a general-purpose language model can also judge candidates point by point, compare pairs, or order a list. those outputs are model judgments too. they can change with prompt or model version, input order, formatting, context limits, and generation instability. an explanation may sound persuasive and still support a bad order. test consistency, output structure, latency, cost, and ranking quality against the same questions. choose between a dedicated reranker and an LLM judgment from measured requirements, not from whether one of them writes a nicer explanation.

06 / EXPANDSearch small. Return enough context

sometimes the right passage is selected and still cannot stand alone.

imagine a policy section as a large parent page. inside it sits a small child passage containing one precise rule. searching the child keeps the match focused. after selection, the system follows the child-to-parent link and returns the containing section, or a bounded neighboring range, so the rule arrives with the definitions and conditions around it.

that is parent-child retrieval: search a small child, return a larger parent when the question needs it. Part 3 owns how those passages were created. this part owns the decision to expand selected evidence.

the link needs stable child and parent identifiers, source version, ordering, permissions, and provenance. if several children point to the same parent, deduplicate it. if several parents could satisfy the expansion, define which one wins. recheck authorization on the expanded object. an allowed child does not make every sibling passage safe to expose.

larger is not automatically better. a parent can restore a missing definition and also pull in unrelated text. duplicated parents consume tokens twice. stale or differently authorized material can cross the boundary. the final context budget includes system instructions, the user’s question, tool text, evidence, and the space left for an answer.

treat that budget as a measured trade-off. compare the child alone, the whole parent, and a bounded neighboring window on the same questions. count the resulting tokens. test ordering, duplication, distractors, truncation, citations, latency, and whether the target generator still uses the evidence. select the smallest context that preserves the facts and dependencies needed for this question.

FIG. 03 — search the focused child; return the needed parent contextPARENT / CHILD

07 / TRACERead the trace before adding a stage

the vacation example gives one bad order, not permission to build every mechanism above.

check the trace in sequence:

  1. eligibility: did a metadata constraint remove the answer before matching began?
  2. candidate recall: after the lexical and representation-based lists were merged, was the answering passage present anywhere?
  3. fusion and ranking: was it present but pushed below passages that only shared the subject?
  4. context sufficiency: did the selected passage need its parent, or did expansion flood the context with noise?

candidate absence and wrong order are different failures. so are a missing parent and an oversized one. preserve the trace, name the failed responsibility, and change one stage. until the evaluation work in Part 8 measures it, that change is a hypothesis.

add the stage that fixes the failure you can observe, then test what it displaced.

one failure remains.

eligibility can be correct. both candidate paths can run. the lists can be merged. the answer-bearing passage can win the rerank. its parent can fit inside the context budget. the final answer can still be wrong because the retrieval task never captured what the person meant.

what if the user’s wording created the failure?

that is where Part 6 starts.

FIG. 04 — five optional jobs and four visible retrieval failuresRETRIEVAL RESPONSIBILITIES

reference notes

  • filtered vector search behavior: pgvector official repository, accessed 21 August 2026. https://github.com/pgvector/pgvector
  • fail-closed authorization guidance: OWASP Authorization Cheat Sheet, accessed 21 August 2026. https://cheatsheetseries.owasp.org/cheatsheets/Authorization_Cheat_Sheet.html
  • heterogeneous lexical and dense retrieval evidence: BEIR paper. https://arxiv.org/html/2104.08663
  • reciprocal rank fusion: Cormack, Clarke, and Buettcher, 2009. https://plg.uwaterloo.ca/~gvcormac/cormacksigir09-rrf.pdf
  • hybrid fusion analysis: Bruch, Gai, and Ingber, 2022. https://arxiv.org/html/2210.11934
  • bi-encoder and cross-encoder roles: Sentence-BERT and Sentence Transformers documentation. https://arxiv.org/html/1908.10084 and https://www.sbert.net/docs/cross_encoder/usage/usage.html
  • LLM ranking limitations: RankGPT and Pairwise Ranking Prompting. https://arxiv.org/html/2304.09542 and https://arxiv.org/html/2306.17563
  • parent-child retrieval guidance: Azure Databricks, accessed 3 September 2026. https://learn.microsoft.com/en-us/azure/databricks/ai-search/retrieval-quality
  • long-context position effects: Lost in the Middle. https://arxiv.org/html/2307.03172
RAG without the magic

Continue reading

View all 8 parts →
FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 5 · RELEVANCE●2026.08.21●FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 5 · RELEVANCE●2026.08.21●