Home / Chapter 4 · Context and knowledge
    Last edited · 8 min read

    Use with AI

    RAG

    RAG (retrieval-augmented generation) means finding passages in your documents and pasting them into the prompt before the question. The model answers from the text it was given instead of from memory, so its knowledge can be current, private and cited, but answer quality depends mostly on what retrieval finds.

    In plain wordsAn open-book exam. Instead of relying on the student’s memory, you put the right page of the textbook in front of them just before the question. Give them the wrong page and they will answer wrongly with the same confidence, because they read what they were given.

    Step through the pipeline. See which chunks drop out at retrieval and reranking, and what changes in the answer

    An employee asks: “How many days of leave do I get here after 10 years?”

    Indexing and chunking

    Retrieval and reranking

    Query-side and index-side tricks

    The prompt and citations

    Evaluation, agents and alternatives

    Check yourself

    How does RAG work, and where does it most often fail?

    RAG gives the model knowledge through its context instead of relying on its memory. Offline, documents are parsed, chunked along their structure, tagged with metadata and permissions, and indexed both as vectors and in BM25. At query time a hybrid search runs with a permission filter, a reranker picks the few best chunks, and the prompt says to answer only from them, with citations, and to admit when the answer is missing. What fails most often is parsing and retrieval, not the model: the right chunk simply is not in the prompt. So measure retrieval recall and answer faithfulness separately.

    Po polsku

    RAG podaje modelowi wiedzę w kontekście zamiast liczyć na jego pamięć. Offline dokumenty się parsuje, tnie po strukturze, opisuje metadanymi i uprawnieniami i indeksuje wektorowo oraz w BM25. Przy pytaniu wyszukiwanie idzie hybrydowo z filtrem uprawnień, reranker wybiera kilka najlepszych fragmentów, a prompt każe odpowiadać tylko z nich, z cytatami, i przyznać, gdy odpowiedzi nie ma. Najczęściej zawodzi parsowanie i wyszukiwanie, nie model: właściwego fragmentu po prostu nie ma w prompcie. Dlatego recall wyszukiwania i wierność odpowiedzi mierz osobno.

    Follow-up questions (5)
    Users report wrong answers. How do you find the cause?
    Go through the traces stage by stage: was the right chunk in the search results, did it survive reranking, did it reach the prompt, and did the model use it? Missing from the results points to parsing, chunking or the query. Present but misused points to the prompt or the model. Stale content points to the indexing process.
    RAG or long context?
    Long context is simpler for a small, stable collection, especially with prompt caching. RAG wins for a large or changing collection on cost per question, latency, permissions and citations, and a shorter context also means less quality loss from length.
    How do you handle permissions?
    ACLs stored in each chunk’s metadata and synced with the source, a filter in the index query based on the user’s identity, and a test that someone without access doesn’t get the chunk. Never through an instruction in the prompt.
    When do you use agentic retrieval instead of a single search?
    When the question needs several steps or a comparison of sources and can’t be answered with one query. Start with a single search, measure which questions it fails on, and pay for the agent’s turns and latency only there.
    When does HyDE hurt?
    When the model doesn’t know the domain: the hypothetical answer contains made-up names, numbers and terms, so its embedding lands next to documents about something else. It also doesn’t help when searching for identifiers and exact phrases, where BM25 wins, or under a tight latency budget. Decide by comparing recall@k with and without HyDE on your own question set.

    Sources

    Report an error · Suggest a fix