FIFTY.DEV
All field notesJournal /How does a search find the same idea in different words?
AI · RAG2026.08.21 · 9 min

How does a search find the same idea in different words?

Embeddings estimate relatedness. Indexes retrieve candidates.

RAG without the magicPart 4 of 8

RAG, without the magic · Part 4 of 8.

01 / VOCABULARYThe words do not match

imagine an HR policy containing this sentence:

“employees are entitled to participate in the automobile coverage program.”

then someone asks:

“what’s the car insurance policy?”

in this hypothetical example, the content-bearing phrases car insurance and automobile coverage program share no terms. the full displayed query and passage do share the stopword the. lexical analyzers may remove stopwords, stem terms, or expand them, so the lexical result depends on the configured analysis rather than proving a universal miss.

that does not make lexical retrieval primitive or obsolete. a lexical retriever can tokenize text, normalize forms, stem words, weight fields, and use configured synonyms. it remains useful when names, policy codes, clauses, dates, or rare terms matter. if its analysis rules include car ↔ automobile and insurance ↔ coverage, it may recover this passage directly.

semantic retrieval adds a different signal. an embedding model can produce representations that let a comparison rank some paraphrases and related wording even when the surface terms do not overlap. here, a suitable setup may place the question near the policy sentence.

“near” needs a boundary. the model estimates that the texts are related under its training and configuration. it does not understand insurance in the human sense. it does not prove that the passage is correct, current, authorized, or sufficient to answer every interpretation of “policy.”

lexical and semantic signals expose different evidence. this part isolates the semantic path so its moving parts are visible. combining those signals belongs to Part 5.

FIG. 01 — two retrieval signals across one vocabulary gapLEXICAL / SEMANTIC

02 / JOBSThree jobs, not one

after Part 3, the source material already exists as searchable passages. semantic candidate retrieval adds three separate jobs.

  1. the embedding model represents text. it converts each stored passage into an ordered list of numbers. when a question arrives, it converts that question with the compatible query setup. these representations are model-specific.
  2. the comparison function orders candidates. it compares the query representation with passage representations using the function and normalization expected by the model and search system.
  3. the index makes lookup practical. it organizes stored representations so the system can find nearby candidates without treating every search as an unstructured scan of the whole collection.

the path is simple enough to say without the machinery:

passage → document representation → index

question → query representation → comparison → ranked passages

the index stores or locates vectors, but the result returned to the rest of the RAG system is still source text with its references. the generator reads that text. it does not need to share the embedding model’s coordinate system.

follow one candidate all the way through: an approved source passage is encoded with the recorded document mode and stored beside its source identity. the question is encoded with the compatible query mode. the declared comparison and index return a ranked candidate identifier. the application then fetches the original passage text and its provenance for context assembly. representations support candidate lookup; source text supports the answer.

an index can perform exact nearest-neighbor search or use an approximate method that examines a more selective set of candidates. that choice changes operating behavior, not the meaning of the representation. either way, the output is a ranked candidate list.

03 / MAPThe map is useful. It is still a sketch

picture a sheet of paper with one marker for the question and several markers for passages. the car-insurance question sits near the automobile-coverage passage. a banana-recipe passage sits farther away. a comparison can now order the markers by their relationship to the query.

that is the whole useful part of the map analogy.

real embeddings are not a tidy two-dimensional sheet. they contain many learned coordinates, and individual axes usually do not have stable names such as “transportation” or “insurance.” a flat drawing does not preserve every neighborhood or distance in the real representation. high-dimensional behavior depends on the learned data distribution, comparison function, normalization, and index configuration.

angle has no universal human-readable meaning. magnitude does not inherently mean more text, more detail, more confidence, or a more complete answer. adding dimensions does not automatically improve retrieval. dimensionality changes the representation shape and affects storage and computation, but quality belongs to the configured model and task, not to a larger coordinate count.

so the map is explanatory compression. nearby markers mean only that this configured system gives them a high comparison score or a small distance. that is an estimate of relatedness, not a certificate of synonymy, truth, or relevance.

FIG. 02 — the map is a teaching aid, not a literal semantic coordinate systemANALOGY LIMIT

04 / COMPARISONThree ways to compare

once text has become vectors, the system needs a rule for ordering them. three common functions appear often.

cosine similarity compares vector orientation after accounting for vector lengths. it does not make orientation equivalent to meaning. it is useful when the model’s contract expects that comparison and the vectors are handled as documented.

dot product multiplies corresponding values and adds the results. without normalization, vector magnitude can affect the ordering. that does not mean dot product prefers longer passages or more comprehensive answers. it has no concept of either.

Euclidean distance measures straight-line separation between vector endpoints. a smaller distance usually means a nearer candidate, although a system may expose a converted score with the direction reversed. Euclidean distance is not inherently more precise. precision is an observed retrieval outcome, not a property granted by the name of a distance function.

normalization can rescale vectors, often to unit length. under unit normalization, cosine, dot product, and Euclidean distance can produce equivalent rankings when implemented consistently. without the same assumptions, they can behave differently.

there is no durable “best for” table here. read the embedding model’s documentation for its intended comparison, query and document formatting, and normalization. then verify that the index uses the same contract and that larger or smaller scores mean what the integration expects. geometry alone cannot choose the right function.

FIG. 03 — comparison functions have mathematical jobs, not prose preferencesNO UNIVERSAL WINNER

05 / COMPATIBILITYCoordinates need a shared system

a coordinate only makes sense inside its coordinate system. latitude and longitude, a street address, and a building’s local floor grid can all describe the same place. the raw values are not directly comparable unless a defined transformation connects them.

embeddings work the same way. passage and query representations must come from an explicitly compatible contract. the common case uses one model and revision for both, with the model’s supported document mode and query mode. another valid setup may use separate query and document encoders that were intentionally trained as a compatible pair. some models require different prefixes or task modes for each side.

“use the exact same call everywhere” is therefore too broad. use the model-specified query and document modes, preprocessing, prefixes, output dimension, normalization, and comparison function.

equal dimensions are only equal shape. two vectors can each contain the same number of values and still belong to unrelated learned coordinate systems. a database may accept both arrays while the comparison is meaningless. this is the quiet failure: storage succeeds, retrieval degrades, and nothing throws an error.

record the representation contract with the stored set: model and revision, query or document mode, preprocessing, dimension, numeric type, normalization, comparison function, and index configuration. if that contract changes and compatibility is not explicitly documented and verified, build a new embedding set and index. compare it before switching over. then retire the old one. re-embedding and reindexing are migrations, not cleanup chores.

input handling belongs to that contract too. embedding inputs are bounded by a model, tokenizer, or API policy. an oversized input may be rejected, truncated under a declared setting, or split by the application. check the selected model’s current behavior and treat truncation as an explicit loss of information, not a hidden default to hope for.

FIG. 04 — equal dimensions do not make unrelated representation spaces compatibleMODEL CONTRACT

06 / REQUIREMENTSChoose from requirements, not rankings

the published version of this article tried to rank models and vector databases. that table aged faster than the lesson.

start with the work instead.

for the embedding model, ask what language, domain terms, identifiers, negation, and query-to-passage task it must handle. check its query and document modes, input policy, privacy path, deployment constraints, throughput, cost, and maintenance burden. most of all, check whether it recovers useful candidates from representative questions over the actual corpus. public benchmarks can narrow a list. they cannot choose for this system.

for storage and indexing, ask about corpus size and growth, insert and delete behavior, filtering, tenancy, consistency, backup and recovery, observability, data location, portability, and who operates it. include the database the team already knows. a specialist vector service is one option, not a requirement. a general database with suitable vector types, operators, and indexes may fit when joins, transactions, or operational consolidation matter.

managed versus self-operated is an ownership boundary, not a maturity ladder. compare staffing, reliability work, privacy, control, recovery, upgrades, and measured total cost. there is no universal traffic threshold that makes the decision for you.

the same rule applies to model size, dimension count, and index family: write down the requirement, configure a candidate correctly, and measure the resulting system on its own material. no product name can skip that work.

a minimum contract test keeps that comparison honest. freeze a small corpus and representative query set. record the model revision, document and query modes, preprocessing, normalization, comparison direction, and index settings. first compute a brute-force exact ordering as a reference. then compare the configured index ordering—and any approximate index under its declared settings—against that reference. change one variable at a time so a gain or regression has an owner.

07 / RELEVANCENearest is not the same as useful

now the search is working. it can recover the automobile-coverage passage for the car-insurance question without claiming that the model understood either phrase. the three jobs are separate. the coordinate contract is explicit. the storage choice follows requirements.

then, in an illustrative leave-policy example, a user asks:

“how do i apply for leave?”

the nearest candidate says:

“annual leave is 21 days.”

a lower candidate says:

“submit form HR-101 to your manager.”

the first passage is plainly related to leave. it is also the wrong evidence for the requested action. the index did its job: it returned what the configured comparison placed nearby. relatedness supplied a plausible candidate, not an answer.

how does the system turn related candidates into evidence that answers the question?

that is where Part 5 starts.

FIG. 05 — choose from the work; nearest can still be wrongRELATED ≠ ANSWER-BEARING

reference notes

  • lexical text analysis and similarity settings: Elasticsearch documentation, accessed 21 August 2026. https://www.elastic.co/docs/manage-data/data-store/text-analysis and https://www.elastic.co/docs/reference/elasticsearch/index-settings/similarity
  • heterogeneous retrieval baseline evidence: BEIR paper. https://arxiv.org/html/2104.08663
  • semantic-search and query/document encoding guidance: Sentence Transformers documentation. https://www.sbert.net/examples/sentence_transformer/applications/semantic-search/README.html
  • comparison functions and normalization relationships: Faiss documentation. https://github.com/facebookresearch/faiss/wiki/MetricType-and-distances
  • exact and approximate vector search in PostgreSQL: pgvector official repository, accessed 21 August 2026. https://github.com/pgvector/pgvector
RAG without the magic

Continue reading

View all 8 parts →
FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 4 · REPRESENTATIONS●2026.08.21●FIFTY.DEV●RAG · WITHOUT THE MAGIC●PART 4 · REPRESENTATIONS●2026.08.21●