Home / Chapter 4 · Context and knowledge
Last edited · 9 min read
Embeddings and vector search
An embedding model turns a whole passage of text into a single vector, so that texts with similar meaning get nearby vectors. Retrieval in RAG rests on this: the question becomes a vector too, and the database returns the passages with the nearest vectors, even when they share no words with it.
In plain wordsA map where every passage of text has its own pin, and texts with similar content sit close together. Searching means pushing in a pin for the question and collecting its nearest neighbours. A contract number barely moves a pin, so contracts 48213 and 48231 sit almost on the same spot.
Pick a question or type your own and see which passage each method puts on top
The whole knowledge base: 10 passages
A simulation, not a real model: an embedding model can’t run here. Each passage has hand-assigned weights for eight named concepts, and the words of the question take their weights from a small dictionary, so words outside it change nothing. A real vector has hundreds of unnamed dimensions. Cosine, BM25 (a simplified stemmer strips common endings) and RRF are computed live. The companies and contract numbers are made up.
One vector for the whole passage
- An embedding model, usually a transformer, computes a vector for every token and combines them into one: by averaging them (mean pooling) or, in embedders built on a decoder LLM such as Qwen3-Embedding (2025), by taking the last token’s vector. The result has a fixed length, typically a few hundred to a few thousand numbers, however long the text. This is not the same as the vectors inside an LLM from “Vectors and matrices”: there every token has its own vector, and the model learns to predict the next token, not to compare texts.
- The model is trained contrastively, on pairs: a question and the passage that answers it should get nearby vectors, while the question and the other passages in the same batch should end up far apart. So “close” means “about the same thing, or answers it”, within the limits of what was in the training data.
- Closeness is measured by the cosine of the angle between two vectors. If the vectors have length 1 (some models return them that way; for the rest you normalise them yourself), cosine is just the dot product, and ranking by Euclidean distance comes out identical.
Searching millions of vectors
- Exact search (brute force) compares the question with every vector. The index takes number of vectors × dimensions × bytes, e.g. 10 million passages × 1024 dimensions × 4 B (float32) ≈ 41 GB, and every query reads all of it. At tens of thousands of passages that means milliseconds and no missed results; at millions, with heavy traffic, it is too slow and too expensive.
- ANN (approximate nearest neighbours) checks only a fraction of the vectors, at the cost of sometimes missing a neighbour. HNSW builds a multi-layer graph of “who is close to whom” and descends it greedily towards the question: fast and accurate, but for low latency the vectors and the graph sit in RAM. DiskANN (2019) keeps only compressed vectors in RAM and the graph with full vectors on SSD: in the paper, a billion vectors on one machine with 64 GB of RAM, under 3 ms per query. IVF splits the vectors into clusters (k-means) and searches only the few nearest ones: less memory, especially with compression, but usually lower recall at the same speed. You tune the trade-off with one parameter,
efSearchin HNSW ornprobein IVF, and measure recall against brute force on a sample of queries. - Vector quantisation cuts memory: int8 takes 4 times less, binary (1 bit per dimension) 32 times less, so about 1.3 GB instead of 41 GB. In a Hugging Face test, binary vectors alone kept about 92.5% of retrieval quality; rescoring the top candidates with the full-precision question vector against the same binary vectors raised that to about 96%, at no extra memory. The figures depend on the model.
- The other lever is fewer dimensions. Models trained with the Matryoshka method keep the most important information in the first dimensions, so a stored vector can be truncated and renormalised without re-embedding the corpus. OpenAI reports that text-embedding-3-large truncated from 3072 to 256 dimensions scores better on the MTEB benchmark than the older ada-002 with 1536.
Weak spots, hybrid search and reranking
- A vector captures the general meaning and loses the details. To a vector, “contract 48213” and “contract 48231” are almost the same thing: a contract with some number. The same goes for codes, numbers, product names the model never saw in training, and company jargon. “Remote work requires approval” and “does not require approval” also get almost the same vector: in the NevIR benchmark (2023) most retrieval models, including the best ones, ranked documents that differ only by a negation no better than chance; cross-encoders did best, and a 2025 reproduction found listwise LLM rerankers better still, though below humans.
- Vector search always returns k results, even when the answer isn’t in the database. A fixed similarity threshold carried over from another model won’t work, because the scale depends on the model: in multilingual-e5, scores cluster between 0.7 and 1.0. You tune the threshold on your own data, and the reranker’s score is a more reliable signal.
- That is why hybrid search is the standard: BM25, the classic ranking by shared words weighted by how rare they are, plus vectors. The two lists are merged with, for example, RRF (reciprocal rank fusion): a document gets the sum of 1/(60 + rank) from each list. RRF looks only at ranks, so there is no need to reconcile BM25 and cosine scales. BM25 needs a stemmer or lemmatisation, otherwise “contract” and “contracts” count as different words; in a heavily inflected language such as Polish this is essential.
- Last comes the reranker, usually a cross-encoder: it reads the question and the passage together, so it sees negation and details that two separately computed vectors can’t capture. It is too slow for the whole corpus, so it only reorders the top candidates, e.g. 50–100 of them. In Anthropic’s test (contextual retrieval, 2024), adding context to passages plus BM25 cut the share of relevant passages missing from the top 20 (1 − recall@20) from 5.7% to 2.9%, and reranking brought it to 1.9%. More in “RAG”.
- Techniques that work around the weaknesses of one vector per passage (HyDE, questions generated for each passage, multi-query, small-to-big, ColBERT-style late interaction) are covered in the “Query-side and index-side tricks” section of “RAG”.
Decisions when building the index
- Chunking decides what a vector means. A long passage averages several topics and matches no question well. A short one loses context: “during this period you are entitled to…”, but which period? Prepending the document and section title before embedding helps. Text longer than the model’s limit is truncated: after 512 tokens in multilingual-e5, after 32k in Qwen3-Embedding.
- A question and a document are different kinds of text: a few words versus a paragraph. Some models expect prefixes, e.g.
query:andpassage:in e5, or an input-type parameter in the API. A missing or swapped prefix throws no error; it silently degrades the ranking. - For a language other than English, such as Polish, choose a multilingual or language-specific model and test it on questions in that language, not just on the MTEB leaderboard. The PIRB benchmark (41 Polish retrieval tasks) compares more than 20 models and shows that a hybrid with keyword search improves even the best vector models.
- Changing the embedding model means re-embedding the whole corpus, because vectors from two models are not comparable, even with the same number of dimensions. Store the source text and the model version with every vector, build the new index next to the old one, and switch when it wins on your test set.
- Evaluate retrieval separately from answers, on a set of questions with the relevant passages labelled. Recall@k: the share of relevant passages found in the top k results. MRR: the mean of 1/rank of the first relevant result, so rank 1 gives 1 and rank 3 gives 0.33. Set k to the number of passages you actually put into the prompt. See “Evals”.
Check yourself
How does vector search work, and why is it not enough on its own?
An embedding model turns a whole passage into one vector, trained so that texts with similar meaning land close together. The query is embedded with the same model, and search returns the nearest vectors by cosine similarity. At millions of passages an approximate index such as HNSW trades a little recall for milliseconds and costs memory, which quantisation reduces. Vectors catch paraphrases but miss identifiers, numbers, rare names and negation, and they always return something. So fuse them with BM25 via RRF, rerank the top candidates with a cross-encoder, and measure recall@k separately from answer quality.
Po polsku
Model embeddingowy zamienia cały fragment w jeden wektor, wytrenowany tak, by teksty o podobnym znaczeniu leżały blisko. Pytanie liczy się tym samym modelem, a wyszukiwanie zwraca najbliższe wektory po cosinusie. Przy milionach fragmentów indeks przybliżony, np. HNSW, oddaje ułamek recall za milisekundy i kosztuje pamięć, którą zmniejsza kwantyzacja. Wektory łapią parafrazy, ale gubią identyfikatory, liczby, rzadkie nazwy i negację, a do tego zawsze coś zwracają. Dlatego łącz je z BM25 przez RRF, czołówkę porządkuj cross-encoderem, a recall@k mierz osobno od jakości odpowiedzi.
Follow-up questions (5)
- You are changing the embedding model. What happens to the index?
- You re-embed the whole corpus, because vectors from different models are not comparable, even with the same number of dimensions. You build the new index next to the old one, compare recall@k on the same question set, and only then switch traffic. That is why you store the source text and the model version with every vector.
- HNSW or IVF?
- HNSW gives high recall at low latency, as long as the vectors and the graph fit in RAM. IVF splits the vectors into clusters and searches a few of them. With PQ compression it fits far more vectors into the same memory, at the cost of recall, which you win back with a larger nprobe or by rescoring the top results on full vectors.
- How do you drop passages when the answer is not in the database?
- You tune a cosine threshold on your own test set, which also includes questions with no answer in the database, because the scale depends on the model. The reranker’s score is a more reliable signal. And the prompt explicitly allows the model to say “this is not in the sources”.
- How do you combine vector search with filters such as permissions or dates?
- A filter applied after the search cuts results out of the top k, so with a narrow filter too few are left, or none. It is better to filter while traversing the index, which many vector databases support, or to keep separate indexes, e.g. one per customer. Enforce permissions at retrieval time, not in the prompt.
- Recall@10 is high, but the answers are still weak. Where do you look?
- First, whether the relevant passage lands at ranks 8–10 while only the top three go into the prompt: a reranker helps there. Next, whether the passage, once cut out of its document, actually contains the answer (chunking). Finally, whether the model distorts it, which is a job for generation evals.