RAG, without the magic · Part 7 of 8.
01 / BASELINEThe pipeline is part of the answer
the fixed pipeline looked fine:
retrieve eligible evidence → assemble context → generate or abstain[1]
then, in an illustrative failure set, four questions reached it.
one was a conversational acknowledgment that needed no retrieval. one compared two independent policies. one retrieved passages that shared words with the question but did not contain the answer. one asked which managers had employees who took parental leave.
the pipeline handled every request the same way. that was the failure.
it was not proof that fixed retrieve-and-generate was obsolete. the same path still handled clear, single-topic questions from a well-matched corpus. if a frozen evaluation shows acceptable retrieval, answer support, abstention, latency, cost, freshness, and access control for the actual query set, the baseline is a valid destination.
a hammer is not broken because one repair needs a screwdriver.
that is the rule for this part: begin with the hammer. keep the failure trace. add the smallest mechanism that addresses that failure. remove it if the evidence does not justify its operating burden.
the labels that follow do not come from one standardized taxonomy. corrective retrieval, Adaptive-RAG, agentic retrieval, and GraphRAG emerged from different papers and practitioner traditions. they overlap. they combine. they are not rungs on a maturity ladder.
02 / CORRECTIVEWhen the evidence looks right and answers nothing
ask about education reimbursement. retrieval returns passages about performance assistance and technical assistance. the vocabulary overlaps. the answer does not.
the fixed path has no independent step that asks whether the evidence is sufficient before generation.
the smallest change worth testing is a generation-time sufficiency instruction:
use the retrieved context only if it supports the answer. otherwise abstain.
call this CRAG-lite as editorial shorthand, not as a paper-defined method. it can help the generator ignore weak context or refuse. it still needs prompt validation and monitoring. it can also refuse when answer-bearing evidence could have been recovered. it does not independently assess evidence, fetch a new source, or repair retrieval.
the named Corrective Retrieval Augmented Generation, or CRAG, paper is more specific. its evaluator maps retrieval confidence to Correct, Incorrect, or Ambiguous. the paper refines retrieved knowledge and uses web evidence for weak or ambiguous retrieval.[2] that external source is part of the paper's setup, not a harmless default. any fallback changes trust, privacy, licensing, freshness, and authorization assumptions.
a broader fuller corrective path can use an explicit assessment gate, then choose one bounded action: use the evidence, rewrite and retry an authorized corpus, switch to another allowed retrieval mode, ask for clarification, use an authorized fallback, abstain, or escalate. query rewrite and same-corpus retry are valid corrective designs. they are not the named CRAG paper's defining mechanism.
every retry needs a reason, a budget, a no-new-evidence check, and a terminal state. stop when evidence is sufficient, candidates repeat, authorization denies the next source, the budget expires, authoritative sources conflict, or escalation is required.
the burden is an assessment step, possible re-retrieval, more state, and another policy to operate. keep the fuller path only if measured results show it recovers answer-bearing evidence that the prompt guard wrongly refused or the baseline missed, without unacceptable new failures or cost. if it mostly repeats candidates, return to the smaller guard.
03 / ADAPTIVEWhen one standing path serves unlike work unevenly
now suppose the trace shows distinct classes of requests need genuinely different pipelines. some clear lookups do well on the fixed path. conversational exchanges need no retrieval. recurring relationship questions benefit from a graph path. open-ended research sometimes needs bounded tool choice.
this is where adaptive architecture becomes useful.
hospital triage is the durable analogy. assess the incoming case once, then send it to a standing specialist path built and tested for that class. triage does not prove the specialist is correct. the router needs an uncertain outcome, a safe fallback, and a record of why it chose the path.
Part 6 routed one query to a treatment such as clarify, rewrite, decompose, or predefined multi-hop. Part 7 draws a different boundary:
Part 6 decides what this query needs. Part 7 decides which durable, governed paths the system is prepared to execute.
the durable router acts before any query-level treatment inside the selected path. each route needs a versioned contract, shared authorization and provenance, independent tests, a circuit breaker, drift monitoring, rollback, and a retirement condition. the named Adaptive-RAG paper used a complexity classifier to choose no retrieval, single-step retrieval, or iterative retrieval in its benchmark setup.[3] the route list above is an illustrative local design, not the paper's universal taxonomy.
maintaining several pipelines plus their router is the burden. a wrong route and fallback can make the system slower or worse. keep durable routing only when measured query slices require materially different paths and the net outcome survives route errors and maintenance. remove routes when they converge, go unused, or lose to one simpler path.
04 / AGENTICWhen the next search cannot be planned in advance
Part 6's decomposition and multi-hop sequences are decided before or by a defined dependency: search for one sourced fact, then use it to form the next search.
another failure looks different. the answer depends on a policy the original question never named, and a fixed plan repeatedly stops too early. the missing dependency appears only after reading the first result.
agentic retrieval addresses that seam. the model participates in a bounded control loop. it chooses from approved tools, observes typed results, and either gathers more evidence or stops under explicit rules.[4][5]
imagine a user asks whether a transfer changes a benefit. the first authorized document-store search returns the transfer policy and mentions that eligibility follows a separate location policy. the model may issue one additional call for that named policy, but only through an allowed tool and only inside the caller's scope.
a trace might read:
request scope: employee-self | region-a
01 search_document_store("transfer benefit eligibility")
result: transfer policy v4 cites location policy v7
02 search_document_store("location policy v7 eligibility")
reason: dependency found in authorized evidence
03 final_answer
stop: information need satisfied
the tool set is finite. inputs and outputs are typed. credentials are user-scoped and minimum-permission. authorization is checked before every call and enforced downstream, outside the model. a call may narrow access, never broaden it. observations are data, not instruction authority.
the loop also needs limits on steps, tokens, time, breadth, and spend. stop on satisfied evidence, clarification required, low confidence, authorization denial, repeated state, an equivalent call with no new information, or exhausted budget. high-impact or irreversible actions require approval where appropriate. log observable calls, sanitized inputs and outputs, evidence identifiers, authorization decisions, model and prompt versions, budget use, and the stop reason. hidden chain-of-thought is not an audit requirement.
no agent framework is required. a small explicit state machine can implement the same contract.
the burden is variable work, more model and tool calls, harder reproduction, and new failure surfaces: wrong tool, malformed arguments, poisoned observations, loops, and unsupported synthesis. keep the agent only if it beats a predefined plan on labeled unknown-path tasks after policy violations, repeats, abstentions, latency, and cost are counted. if a fixed plan catches up, retire the loop.
05 / GRAPHWhen the relationship keeps being rediscovered
consider the question:
Which managers have employees who took parental leave last year?
the answer depends on an explicit chain:
manager → manages → employee → took → parental leave
vector or lexical retrieval can find passages about managers, employees, and leave. it can supply candidates, support multi-hop retrieval, or locate an entry point into a graph. it does not directly store that explicit relationship as a reusable edge.
Part 6's predefined multi-hop can still answer the question. it may first identify employees with the leave record, then retrieve their managers. the same relationship is reconstructed on every similar request.
graph retrieval moves selected relationship work into a maintained structure. entities such as people, policies, teams, and leave records become nodes. directed, typed relationships connect them. a direct structured query can follow a known relation. a local graph-assisted path can start from semantically related entities, expand connected records, bring in associated source text, and generate from a bounded context. a global path can generate over prebuilt community reports to answer corpus-wide questions. vector-plus-graph retrieval can find a flexible entry point, expand explicit relations, then return to raw text for source detail.[9][10]
those are possibilities, not one universal GraphRAG menu. Microsoft GraphRAG is a specific method and library. its primary paper builds an entity graph, creates hierarchical communities and summaries, then uses map-reduce generation over those reports for global questions.[6] current documentation also distinguishes local, global, and basic vector search.[7] global search is resource-intensive.[8] GraphRAG is not a synonym for one cheap traversal.
the graph itself begins with design choices. which questions must it answer. which entity and relation types matter. which direction and time semantics apply. which sources win when they conflict. how provenance and authorization partitions attach to every node and edge.
extraction is not resolution. extracting John Smith creates a mention and candidate attributes. entity resolution asks: is this John Smith the same John Smith already in the graph? a false merge joins two people. a false split creates two records for one person. relationship extraction then proposes an edge. provenance ties that assertion to source text, document and version, effective time, extraction version, confidence, and review state.
extraction is use-case-specific. the sentence John Smith approved the Project Phoenix budget in Conference Room B can support a project-approval graph, a finance graph, or a facilities graph. extracting every noun is not a schema.
freshness is the quiet cost. a new, changed, deleted, or permission-changed document may require affected text to be reprocessed, entities to be re-resolved, edges to be added or retracted, communities and summaries to be refreshed, embeddings and caches to be invalidated, and deletion to be verified across derived stores. incremental updates are possible, but correctness is implementation-specific.[11] vector systems have versioning, deletion, access-label, provenance, and cache burdens too. the graph adds more derived state to reconcile.
keep graph retrieval only when measured relationship or corpus-wide questions improve enough to justify extraction, resolution, provenance, search-mode, and freshness work. keep purely content-based questions on simpler retrieval. remove the graph path if false merges, stale edges, or lifecycle cost erase the gain.
06 / DECISIONOne flat decision, not a ladder
| pattern | measured failure | added mechanism | operational burden | evidence to keep it | |---|---|---|---|---| | fixed baseline | no target failure observed | none | baseline retrieval, generation, provenance, freshness, and access control | representative evaluation remains acceptable | | corrective | evidence overlaps in words but not in answer support | generation guard, then explicit assessment and bounded repair if needed | assessment, retry state, source policy, possible re-retrieval | fewer unsupported answers or more recovered evidence without unacceptable false refusals or cost | | adaptive | distinct query slices need different standing paths | versioned router with uncertainty, fallback, and durable specialized paths | multiple pipelines, route labels, drift, rollback, debugging | per-slice gain survives route errors and maintenance | | agentic | the evidence path is unknown in advance but recoverable through approved tools | bounded model-controlled tool loop | variable calls, tool policy, loops, reproducibility | wins on labeled unknown-path tasks within policy and resource limits | | graph | recurring questions depend on explicit reusable entities and relationships, or corpus-wide structure | maintained graph with direct, local, global, or vector-plus-graph retrieval as needed | extraction, resolution, provenance, updates, deletions, summaries, mixed retrieval | target relationship or global slice improves enough to cover lifecycle cost |
the rows can combine. an adaptive router can select a graph-assisted backend. a bounded agent can receive both document-search and graph-query tools. composition does not remove the burden of either mechanism.
none of these pattern names creates safety. high-stakes use still needs authorization enforced on every request and downstream action, citation and evidence validation, effective-date and version controls, calibrated abstention, human review where consequences warrant it, tool approval and downstream controls, monitoring, incident response, and evaluation.[13][14]
the architecture is only one hypothesis in that control system.
Part 8 takes the next question: what evidence would prove that the added path fixed its target failure without creating a worse one?
architecture is still a hypothesis until its failure rate and operating burden are measured.
reference notes
[1] Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks: https://arxiv.org/abs/2005.11401
[2] Corrective Retrieval Augmented Generation: https://arxiv.org/abs/2401.15884
[3] Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity: https://arxiv.org/abs/2403.14403
[4] ReAct: Synergizing Reasoning and Acting in Language Models: https://arxiv.org/abs/2210.03629
[5] Toolformer: Language Models Can Teach Themselves to Use Tools: https://arxiv.org/abs/2302.04761
[6] From Local to Global: A Graph RAG Approach to Query-Focused Summarization: https://arxiv.org/abs/2404.16130
[7] Microsoft GraphRAG query overview: https://github.com/microsoft/graphrag/blob/7bb23cc7f32f47cf618a1ae9cca39a6695f434ae/docs/query/overview.md
[8] Microsoft GraphRAG global search: https://github.com/microsoft/graphrag/blob/7bb23cc7f32f47cf618a1ae9cca39a6695f434ae/docs/query/global_search.md
[9] Microsoft GraphRAG local search: https://github.com/microsoft/graphrag/blob/7bb23cc7f32f47cf618a1ae9cca39a6695f434ae/docs/query/local_search.md
[10] Microsoft GraphRAG default dataflow: https://github.com/microsoft/graphrag/blob/7bb23cc7f32f47cf618a1ae9cca39a6695f434ae/docs/index/default_dataflow.md
[11] Microsoft GraphRAG incremental index: https://github.com/microsoft/graphrag/blob/7bb23cc7f32f47cf618a1ae9cca39a6695f434ae/packages/graphrag/graphrag/index/update/incremental_index.py
[13] OWASP Authorization Cheat Sheet: https://github.com/OWASP/CheatSheetSeries/blob/6b8819da79e0537d072e04296ffa3adfc94ba881/cheatsheets/Authorization_Cheat_Sheet.md
[14] OWASP LLM06:2025 Excessive Agency: https://genai.owasp.org/llmrisk/llm062025-excessive-agency/