RAG, without the magic · Part 3 of 8.
01 / FAILUREThe doctor's note is in the document
ask a policy assistant a simple question:
do i need a doctor's note for sick leave?
the source contains the answer. one section says how much sick leave an employee receives. a later sentence says that an absence beyond a stated duration requires a medical certificate.
now put a boundary between those statements.
the retrieved passage contains the entitlement, because it includes the words “sick leave.” the medical-certificate condition sits in another passage, separated from the heading and most of the language in the question. the system returns the first passage and answers with the entitlement, then says the supplied evidence contains no rule about a doctor's note.
the broken version might look like this:
passage one
sick leave policy. employees receive an annual sick-leave entitlement. absences beyond the stated
passage two
duration require a medical certificate. unused entitlement does not carry into the next period.
the first passage has the heading and the obvious match for “sick leave.” it also ends before the condition becomes interpretable. the second passage has the condition, but “duration” has lost the thing it measures and the heading is gone. either passage alone is weak evidence for the question.
a better boundary would keep the complete medical-certificate statement with the sick-leave heading or repeat enough governing context to make the statement stand on its own. the important point is not that every policy needs one large passage. it is that a rule and the context required to read it should not be separated by an arbitrary count.
nothing was missing from the source. the boundary made the answer hard to retrieve as one complete unit.
this is chunking: cutting already usable source material into searchable passages before indexing it. each passage is usually called a chunk. the name is small. the decision is not.
a chunk boundary decides which headings, facts, exceptions, and neighboring steps can be selected and scored together during the first retrieval stage. later stages may retrieve several passages, but the initial units still shape what can travel together.
02 / QUERIESDesign for the questions
chunking is closer to database-schema design than document housekeeping.
a database schema is not designed by asking only, “how can i store this data?” it is designed around the queries the system needs to answer. storage shape affects every later query. changing that shape usually means migrating derived data.
searchable passages work the same way. do not ask only, “how many words fit?” ask:
- what questions will people bring here?
- which facts must stay together to answer them?
- which headings, definitions, exceptions, or steps make those facts interpretable?
- what unrelated material would make the answer harder to isolate?
a useful passage tends to have three properties:
- it is self-contained enough to interpret.
- it stays coherent around one subject or authored unit.
- it is focused enough that the answer is not buried.
those properties pull against one another. a narrower passage can isolate a fact and detach its qualifier. a broader passage can preserve the qualifier and bring several unrelated sections with it. there is no boundary rule that resolves this tension for every document and every question.
there is also a hard practical caveat. systems that turn passages into searchable representations accept limited amounts of text at once. those limits are counted according to a particular tokenizer: the component that divides text into the units a model counts. a token is not a universal word or character. counts change with the tokenizer, model encoding, language, punctuation, code, and the text itself.
check the target service and count with its tokenizer. oversized input may be rejected, truncated under a configured policy, or require splitting before submission. do not assume silent truncation is safe. exact model limits and vendor comparisons belong in a dated implementation note, not a durable boundary rule.
size is therefore a hypothesis constrained by two things: the relationships a question needs and the limits of the actual system.
03 / SPECTRUMFour sources of boundary evidence
four useful boundary families are size-based, structure-aware, topic-aware, and LLM-assisted. they overlap. they are not an exhaustive taxonomy, and they do not form a quality ladder.
size-based splitting
size-based splitting cuts after a chosen count of characters, bytes, words, sentences, or tokens.
think of cutting a cake with a ruler. each slice reaches the chosen width whether the cut lands in sponge, filling, or icing. the result is bounded and predictable. that can be useful for a quick baseline or for records that are already uniform and independent.
the likely failure is equally plain: the count knows nothing about a sentence, a heading, a table, or a relationship between two statements. it may put the question cue on one side and the answer on the other. the chosen counting unit also matters. a character budget and a token budget do not produce interchangeable boundaries across languages or formats.
imagine a directory of short, independent product records. one record per passage may already satisfy the question shape, and a hard size limit can catch the occasional oversized description. now imagine applying the same count to a policy manual. a cut can land between “requires approval when” and the condition that follows. the mechanism is identical. the source structure changes whether the result is useful.
size-based splitting is therefore neither a toy nor a production rule. it is a measurable boundary choice. its simplest test is to place expected questions near proposed edges and inspect whether the complete answer survives.
structure-aware splitting
structure-aware splitting prefers boundaries already present in the document: titles, sections, paragraphs, sentences, clauses, list items, table elements, code symbols, or conversation turns. if one authored unit is too large, a fallback divides it at the strongest reliable internal boundary available.
this is the cake-layers version. follow the layers before reaching for the ruler.
it is a defensible baseline when the parser has recovered reliable structure. it can preserve the author's organization without pretending that every section has the right size. its failure starts upstream: missing headings, flattened tables, malformed extraction, or oversized sections can leave weak structure to follow. a newline-based recursive splitter is a useful fallback procedure, but it is not the same thing as understanding a parsed PDF, HTML tree, or code syntax tree.
take a handbook with a heading called “parental leave,” followed by separate subheadings for eligibility, application, approvals, and benefits. section boundaries provide useful evidence. a structure-aware rule can keep “application” with its numbered steps, then decide what to do when that section exceeds the system's actual limit.
the likely failure is trusting decoration that is not dependable structure. a PDF extractor may turn every line into a paragraph. a repeated page header may look like a new section. a long section may contain several subjects. prefer the strongest reliable element, but keep a measured fallback for elements that are malformed or too large.
topic-aware splitting
topic-aware methods estimate where the subject changes. an editor might separate a company-history paragraph from a company-mission paragraph even when both fit beneath one heading. textual signals can suggest the same break.
those signals may come from repeated vocabulary, changes between adjacent passages, embeddings, classifiers, or a combination. they do not necessarily require embedding every sentence. they also do not guarantee a better boundary. gradual topic shifts can be missed. abrupt wording changes can split one subject. thresholds can create tiny fragments or oversized passages. useful hierarchy can disappear if topic signals ignore the document's structure.
consider free-form meeting notes that move from a release date to a hiring constraint without a heading. a topic signal may detect that change and avoid placing both discussions in one passage. the same signal may split a single incident report when the vocabulary changes from symptoms to diagnosis. the subject did not change, but the words did.
topic-aware selection is useful when topic-boundary failures are visible in real questions. it is not a substitute for preserving speaker labels, section paths, or other structure the source already provides.
LLM-assisted boundary selection
a language model can propose logical boundaries among candidate paragraphs or parsed elements. it may also suggest labels for the resulting passages. the controlled indexing process still owns the decision.
constrain proposals to existing element identifiers or source offsets. validate ordering, gaps, overlaps, duplicates, and out-of-range spans. record the model, prompt, decoding settings, and preprocessing version. keep a deterministic fallback.
this family can help with irregular documents where cheaper rules keep failing. it also adds cost, latency, variability, and validation work. those are measurements for the target corpus, not a reason to place it above the other families.
suppose a long narrative case file contains dated events, quoted correspondence, and commentary with inconsistent headings. an LLM can be asked to choose among known paragraph boundaries and identify the first paragraph of each new episode. the result is still a proposal. if it skips a paragraph, reverses two offsets, or returns overlapping spans, the pipeline rejects it rather than indexing invented structure.
the likely failure is not only an odd boundary. model or prompt changes can alter the proposal on the next rebuild. recording the configuration makes that difference inspectable. constraining choices to source offsets keeps the model from rewriting the document while pretending to segment it.
combinations are normal. structure can define safe outer units while a topic signal refines one oversized prose section. that is a design option, not an inevitable advanced stage.
04 / SIZESize trades focus for retained relationships
consider one hypothetical parental-leave passage:
submit the request form at least 30 days before leave. if an emergency makes that notice impossible, submit it as soon as practical. manager approval is required before scheduling begins.
the question is:
how do i apply for parental leave if an emergency prevents 30 days' notice?
hold that source text and question constant while comparing three boundary choices. these are illustrative outcomes, not token recommendations.
size-based cut through the qualifier. one passage ends after the 30-day rule; the next begins with the emergency exception. retrieving only the first passage supports an incomplete answer, while retrieving only the second can orphan what the exception qualifies.
structure-aware procedure unit. one passage keeps the request rule, emergency exception, and approval step together because they belong to the same application procedure. the unit preserves the relationship the question needs.
oversized mixed-topic unit. one passage keeps the procedure intact but also absorbs eligibility, entitlement, and benefits sections. the needed relationship survives, yet it competes with unrelated policy material and consumes more context.
smaller is not automatically more precise. larger is not automatically safer. choose a candidate large enough to retain the relationships your questions need and small enough to keep the evidence focused and within every downstream limit. then test it.
if numerical candidates are useful in an implementation, count them with the target tokenizer and label them as candidate configurations for that corpus. a number without that setup is decoration.
05 / OVERLAPOverlap is insurance, not a cure
sometimes a sensible boundary still lands beside a needed qualification.
imagine this policy statement:
vacation requests require advance notice. late requests may be denied. manager approval is required for requests beyond the stated duration.
one passage ends after the late-request condition. the next begins with the manager-approval rule. a question about a long vacation may need both.
overlap repeats a limited edge of the first passage in the next one. if the repeated edge includes the notice condition, the second passage can carry that condition beside the approval rule. this can reduce some misses caused by an arbitrary edge.
the copied text is not free. it increases indexed material and search work. near-duplicate passages can crowd out distinct evidence in a limited candidate set. too much overlap can turn several passages into slightly different copies of the same material.
start at zero when boundaries already preserve complete units. add the smallest measured overlap that repairs a demonstrated boundary miss. inspect the actual repeated text, not only a percentage.
overlap cannot repair the wrong unit. duplicating half a sentence leaves half a sentence. repeating rows without their table header leaves orphaned values. copying the tail of a function does not restore its signature. carrying a topic change into the next passage may only mix both topics twice.
insurance helps at the edge. it does not redesign the building.
06 / STRUCTURESome content has its own boundaries
text is not one undifferentiated string. tables, code, questions and answers, lists, clauses, and conversations carry relationships that whitespace alone does not express.
tables. preserve the table title, column headers, units, footnotes, and the row and column relationships needed to read a value. a small table may stay whole. an oversized table may split at a stable row or column group, but each derived passage needs the necessary headers, units, table identity, and source location. converting a table to prose is a transformation. keep the original and check that values and relationships survived.
an illustrative cell containing “42” is not evidence by itself. it may mean 42 days, 42 dollars, or row 42. the header, unit, row label, and footnote can all be part of the answer.
code. prefer language-aware units such as files, modules, classes, functions, methods, declarations, or notebook cells. keep the repository path, language, symbol name, signature, containing scope, relevant explanation, and source offsets. if a unit is too large, split at valid nested syntax nodes or explicit line ranges. a fenced code block marks formatting. it does not prove that everything inside is one useful unit.
a function body without its signature can lose the input types a question asks about. a code example without the paragraph that explains its configuration may be executable and still be useless for the reader's problem.
question and answer pairs. keep the question, answer, category, and source identity together unless one side is independently meaningful and explicitly linked. an answer without its question often loses its subject. a question without its answer is mostly a promise.
headers and lists. keep the section path or lead-in with the items and qualifications it governs. if a list must split, repeat the governing header and preserve numbering and nesting.
“passport, visa, ticket” changes meaning when the missing lead-in was “documents required for international travel.” the list items did not change. their governing statement vanished.
clauses and policies. prefer clause and subclause boundaries while retaining definitions, exceptions, cross-references, hierarchy, and effective date. a clause is not automatically self-contained, and a neat boundary does not make legal retrieval safe.
transcripts and conversations. preserve speaker labels, turn order, timestamps when available, and conversation identity. a single message may depend on an earlier question, correction, or pronoun. test short turn windows or topic segments without discarding the turn structure.
these are unit-preserving patterns, not document-type presets. first recover the structure. then decide which relationships the expected questions require.
07 / EXPERIMENTRun a small boundary experiment
before indexing the full collection, make the first decision small enough to reverse.
- choose one representative document type and freeze its parsed source.
- write only two or three real questions before changing the boundaries. at least one must be multi-part. for example: “what is the sick-leave entitlement, and when is a doctor's note required?” use that question as the cross-boundary or qualification case; use the remaining question or questions for an ordinary fact or another heading, exception, or qualification case.
- mark the complete source span each answer needs.
- build two candidate boundary configurations. hold the source version, parser output, retrieval setup, query handling, candidate cutoff, and answer setup constant.
- inspect whether the selected passage contains the complete answer and the context required to interpret it. note unrelated text and duplicate passages that arrived with it.
- change one boundary variable and run the same questions again.
if configuration a keeps the doctor's-note condition with its heading and configuration b strands it, that is useful evidence for this policy and question set. it is not proof that configuration a wins for support tickets, source code, or transcripts. the experiment narrows one decision without turning it into a slogan.
keep a short decision log:
- document type
- boundary rule
- illustrative size and overlap settings, counted in the actual unit
- question
- expected source span
- selected passage
- missing context, unrelated text, or duplication
- next change to try
record parser failures, over-limit inputs, and answer misuse separately. a boundary change cannot repair a parser that lost the heading, and it should not receive credit or blame for a later answer-stage failure.
this is not the formal evaluation system. Part 8 will handle metrics, datasets, monitoring, and regressions. for now, the useful conclusion is bounded: this rule kept the answer together for this corpus, these questions, and this retrieval setup.
changing the rule changes passage text, identifiers, offsets, and derived representations. version the segmentation configuration. compare in a separate index or namespace where practical. once a rule is chosen, rebuild the affected derived index and verify that old passages are no longer eligible. Part 2's source-version and deletion controls still apply. they do not need another tour here.
the boundary is now explicit, testable, and still provisional.
next comes a different question: once the right passage exists, how can a differently worded question find it?
reference notes
external sources are listed here for verification. they are not calls to action.
- OpenAI Help Center, “What are tokens and how to count them?” https://help.openai.com/en/articles/4936856-what-are-tokens-and-how-to-count-them
- OpenAI Cookbook, “Embedding texts longer than the maximum context length.” https://cookbook.openai.com/examples/embedding_long_inputs
- Cohere, “Embed API v2.” https://docs.cohere.com/reference/embed
- Unstructured, “Chunking.” https://docs.unstructured.io/open-source/core-functionality/chunking
- Marti A. Hearst, “TextTiling: Segmenting Text into Multi-paragraph Subtopic Passages.” https://aclanthology.org/J97-1003
- LumberChunker, “Long-Form Narrative Document Segmentation.” https://arxiv.org/html/2406.17526
- Tree-sitter, “Introduction.” https://tree-sitter.github.io/tree-sitter