FIFTY.DEV
All field notesJournal /Should this be RAG?
AI · RAG2026.08.20 · 10 min

Should this be RAG?

Decide whether the missing piece is knowledge, behavior, or both.

RAG without the magicPart 1 of 8

RAG, without the magic · Part 1 of 8.

“we need company knowledge in the model” sounds like a technical requirement. it is usually three decisions wearing one coat.

does the application need access to information that changes? does the model need to behave more consistently on a repeatable task? or are both problems present?

that distinction comes before embeddings, vector databases, model providers, or any other architecture. if the failure is misdiagnosed, the system can become more elaborate without becoming more useful.

take one illustrative support question: “does the current policy cover this request?” without authorized access to the current policy, a model may answer from older or general patterns, decline, or produce an unsupported answer. when the current policy and its source identity are supplied with the same question, the application can ask for an answer grounded in that evidence. retrieval changes what the model can inspect at answer time; it does not by itself guarantee that the model will follow the evidence or produce the required structure.

this is the first decision in this series: when should an application retrieve information, when should it rely on prompting or fine-tuning, and when should it combine them?

01 / DECISIONThe closed-book test

start with an exam.

a closed-book exam asks the student to learn a pattern before the test and answer from what they retained. an open-book exam lets the student look up relevant material while answering.

fine-tuning is closer to the closed-book side. supervised fine-tuning updates a model’s weights using example inputs and desired outputs. the model becomes more likely to perform the demonstrated task or produce the demonstrated kind of output.

retrieval is closer to the open-book side. retrieval-augmented generation, or rag, searches for candidate source material and supplies selected material to a language model at question time. the model does not need every fact encoded in its weights. it receives relevant material with the question.

the analogy is useful because it separates two needs:

  • knowledge access: what material must be available while the answer is being produced?
  • behavior shaping: how should the model use an input and produce an output repeatedly?

neither side is automatically better. an exam can test memory, research, or both. an application can have the same mix.

FIG. 01 — open book, closed book: retrieval supplies evidence; fine-tuning shapes behaviorKNOWLEDGE / BEHAVIOR

02 / RETRIEVALWhat retrieval changes

a large language model, or llm, generates text from its prompt and the patterns represented in its parameters. those parameters reflect training. they do not automatically contain later information, and they do not give the application authorized access to private records.

without an authorized data connection, a model cannot reliably answer from a company’s policies, product documents, customer records, or current operating data. it may say it does not know. it may answer from general patterns. it can also produce plausible text that the source material does not support.

retrieval changes the material available at question time. a typical rag path does three things:

  1. search an approved collection for material related to the question.
  2. select some of that material and place it in the model’s context.
  3. ask the model to answer using the question and the selected evidence.

rag is a pattern, not one mandatory product stack. the search might use keywords, semantic similarity, structured filters, or several methods together. the source might be a document collection, a database, or a service reached through a controlled tool. what matters at this decision point is that external evidence is selected at question time. the mechanics matter later, but they do not change the diagnosis.

this is useful when answers depend on changing documents, private sources, or inspectable evidence. updating the corpus usually does not require retraining the generator.

“usually” is doing work there.

a changed document is not available merely because someone saved it, and retrieved material is not automatically correct, authorized, sufficient, or used faithfully. freshness, provenance, access control, citation support, cost, and the split between retrieval failure and generation failure remain engineering work. Part 2 follows those mechanics end to end; here, they are boundaries on the decision.

03 / BEHAVIORWhat prompting and fine-tuning change

if the correct evidence is already in the prompt but the output is inconsistent, adding another search system may not address the failure.

start with the smallest intervention. clearer instructions, a better output schema, and a few good examples can often improve style, format, tool use, and task behavior. “often” is the boundary. prompting is not guaranteed to be enough.

fine-tuning becomes worth testing when representative prompts and examples still do not produce sufficiently consistent behavior, or when a smaller model or shorter prompt may meet the same quality target with better measured latency or total cost.

it can help with a narrow classification task, a repeated structured output, a particular tool-calling pattern, or a specialized style. it can also regress performance outside the target task. compare the tuned model with the base model on target examples and guardrail evaluations.

fine-tuning changes learned behavior. it is not a dependable document lookup system. facts encoded only through a training run do not update automatically, and a generated statement cannot normally be traced from model weights back to one training record.

it also does not make generation deterministic or factually guaranteed. the defensible claim is narrower: evaluations may show better task performance, consistency, or efficiency for a defined workload.

04 / COMPOSITEA support-documentation decision

composite example: the following support scenario combines common requirements seen in documentation-heavy saas teams. it is illustrative, not a report of one named deployment.

imagine a support team answering questions across product documentation, troubleshooting guides, api references, known-issue records, and internal runbooks. the sources change as products and procedures change. an agent needs to verify an answer before sending it and, when appropriate, share the supporting document.

the first question is not “which technology wins?” it is “what is missing when this answer is produced?”

ask two questions:

  1. does the model have the evidence it needs at answer time? if not, test retrieval or a purpose-built data connection. the source still has to be authorized, ingested, indexed, retrievable, and evaluated.
  2. when the evidence is present, does the model use it consistently? if not, improve the prompt and examples first. test fine-tuning only if evaluations justify another training and maintenance cycle.

the answers give three starting points:

  • missing evidence → retrieval or a purpose-built data connection.
  • inconsistent use of available evidence → prompting first, then possibly fine-tuning.
  • both problems → combine the approaches and test retrieval and generation separately as well as end to end.

latency, cost, and whether a smaller model can handle the task are evaluation targets. they help compare interventions; they are not a separate diagnosis.

if the workflow requires strict factual or regulatory guarantees, neither retrieval nor fine-tuning is sufficient alone. add source governance, authorization, validation, abstention or escalation, human review, and risk-appropriate testing.

FIG. 02 — two questions separate missing evidence from inconsistent behavior

for this scenario, retrieval is a strong first hypothesis because the answers depend on changing, inspectable sources. a source update can enter the answer path after ingestion and indexing complete, without another generator training run. logged retrieval results can also show which document version and passages were available for an answer.

FIG. 03 — a source update is not searchable until the ingestion path completesUPDATE PATHS

that does not make rag the universal winner. the team still has to test whether the right passages are found, whether access controls hold, whether citations support the claims, whether latency is acceptable, and whether the model uses the evidence correctly.

if the retrieved evidence is useful but the generator keeps producing the wrong structure or mishandling the same task, prompting may be enough. if it is not, a fine-tuned generator may improve measured consistency. the system can be open-book and better trained.

05 / BOUNDARIESWhere the rule breaks

retrieval and fine-tuning are not the only two switches.

sometimes the answer should come from a tool rather than a document search. a price, account balance, inventory level, or permission check may belong behind a typed api call with validation, not inside generated prose.

sometimes the source material is small and stable enough to place directly in the prompt. that is still question-time context, without a full retrieval pipeline.

sometimes prompting handles the required behavior. another training lifecycle would add work without a measured gain.

sometimes retrieval supplies the right evidence but the model uses it poorly. then the application has both a context problem and a behavior problem. combine retrieval with prompt changes or fine-tuning, using examples that contain the same kind of retrieved context the production model will see.

high-stakes workflows need another layer of decisions. retrieval and fine-tuning are model techniques, not complete safety controls. authorization, source governance, evaluation, abstention, escalation, and human review still sit outside them.

06 / DECISIONThe three-way decision

the useful question is not “rag or fine-tuning?”

ask what is missing:

  • if the application lacks changing, private, or inspectable knowledge at question time, test retrieval or a purpose-built data connection.
  • if the evidence is already present but task behavior is inconsistent, improve the prompt and examples, then test fine-tuning when evaluations justify it.
  • if both constraints exist, combine the approaches and test each stage as well as the whole system.

begin with the smallest intervention that targets the observed failure. measure it against a baseline. keep the boundaries visible.

part 2 moves from the decision to the path: how material is ingested, indexed, retrieved, and placed in front of the model. that is where the open-book analogy becomes a system.

reference notes

  1. OpenAI, optimizing llm accuracy, accessed 20 august 2026. provider guidance on context failures, behavior failures, evaluation, and combining retrieval with fine-tuning.
  2. OpenAI, supervised fine-tuning, accessed 20 august 2026. provider documentation on training examples, weight updates, task performance, and distillation.
  3. Microsoft Learn, retrieval augmented generation and indexes, updated 20 may 2026. provider guidance on retrieval, citations, access control, latency, cost, and hallucination despite grounding.
  4. aws, sync your data with an Amazon Bedrock knowledge base, accessed 20 august 2026. provider documentation on ingestion, parsing, chunking, embedding, and reindexing after source changes.
  5. Lewis et al., retrieval-augmented generation for knowledge-intensive nlp tasks, neurips 2020. the historical research reference for retrieval-augmented generation, parametric and non-parametric memory, and provenance motivation.
RAG without the magic

Continue reading

View all 8 parts →
FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 1 · THE DECISION●2026.08.20●FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 1 · THE DECISION●2026.08.20●