CertSafari
    CCAR-P · Lessons

    Domain 3 · Lesson 16/38

    Contextual Retrieval and Reranking for RAG Indexes

    Design a RAG pipeline with appropriate chunking and indexing strategies

    8 min read
    2.38% of exam
    4 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Explain why isolated chunks retrieve poorly, and how contextual embeddings and contextual BM25 fix that at indexing time
    • Estimate the cost of contextualizing a corpus, and explain the role prompt caching plays
    • Add a reranking stage and choose how many chunks to pass to the model
    • Evaluate retrieval with Pass@k or recall@k, and weigh the alternative indexing ideas the sources compare

    1.Giving every chunk its missing context

    Standard RAG splits documents into chunks of a few hundred tokens before indexing them. That works until a chunk depends on context it no longer carries. Anthropic's example is a set of SEC filings queried with "What was the revenue growth for ACME Corp in Q2 2023?" The right chunk reads: "The company's revenue grew by 3% over the previous quarter." On its own, that chunk doesn't say which company or which period it describes. Neither the embedding index nor the keyword index has anything linking it to ACME or Q2 2023.

    Contextual Retrieval fixes this during preprocessing. It prepends a short, chunk-specific explanation to each chunk before the chunk is embedded (Contextual Embeddings) and before it goes into the BM25 index (Contextual BM25). Writing that context by hand for millions of chunks isn't practical, so Claude writes it. The post used this prompt with Claude 3 Haiku:

    The contextualizer prompt from Anthropic's Contextual Retrieval post: the whole document plus one chunk goes in, a short situating context comes outtext
    <document>
    {{WHOLE_DOCUMENT}}
    </document>
    Here is the chunk we want to situate within the whole document
    <chunk>
    {{CHUNK_CONTENT}}
    </chunk>
    Please give a short succinct context to situate this chunk within the overall document for the purposes of improving search retrieval of the chunk. Answer only with the succinct context and nothing else.

    The generated context is usually 50 to 100 tokens. Sending the whole document once for every chunk sounds expensive, and prompt caching is what makes it affordable: the document is cached once and each chunk's request reads from that cache. With 800-token chunks and 8k-token documents, the post puts the one-time cost at $1.02 per million document tokens. The generic prompt works well, but a domain-specific one can do better, for example one that includes a glossary of terms defined elsewhere in the corpus.

    The embedding layer offers a second route. Claude's embeddings guide lists Voyage's contextualized chunk embedding models, voyage-context-4 and voyage-context-3, with a 120,000-token context length. You call them with contextualized_embed() instead of embed(). They produce chunk vectors that already reflect the full document, with no manual metadata augmentation.

    A support-bot pipeline retrieves the same set of product documentation chunks for every user session and appends them to the system prompt, followed by whatever question the current user asks. The team wants to use prompt caching to cut latency and cost across the many queries that reuse this documentation. Where should they place the cache breakpoint?

    Sources12

    2.Reranking and deciding how many chunks to pass

    The indexing changes improve recall. A reranking stage improves the order of what comes back. A reranker takes the query and the candidate chunks and returns them sorted by relevance, so the chunks most worth including rise to the top before you cut the list down. Claude's embeddings guide lists two Voyage rerankers, both called with rerank(): rerank-2.5 for highest accuracy, recommended for most applications, and rerank-2.5-lite, optimized for latency and cost.

    Top-20-chunk retrieval failure rate in Anthropic's Contextual Retrieval experiments (1 minus recall@20)
    ConfigurationFailure rateReduction vs. baseline
    Standard embeddings (baseline)5.7%—
    Contextual Embeddings3.7%35%
    Contextual Embeddings + Contextual BM252.9%49%
    Contextual Embeddings + Contextual BM25 + rerankingNot stated as an absolute rate67%

    The last knob is K, the number of chunks you put in the prompt. More chunks raise the odds that the relevant one is included, but the post warns that extra material can distract the model. Of the 5, 10 and 20 chunks it tested, 20 performed best, and it recommends experimenting on your own use case rather than treating 20 as a rule. The post also found that contextualizing helped with every embedding model it tested, with Gemini and Voyage embeddings especially effective.

    A platform team is redesigning their RAG pipeline after an audit found a high rate of retrieval failures, where relevant chunks exist in the knowledge base but are not returned for matching queries. They can implement several changes before the next release. Which of the following changes are documented to measurably reduce retrieval failure rates? (Select all that apply.)(Select 3)

    Sources21

    3.Measuring retrieval and weighing alternative indexing ideas

    Every chunking, indexing or reranking change should be measured against a baseline. The contextual embeddings cookbook builds an evaluation set of 248 queries, each paired with a 'golden chunk', and scores Pass@k: whether that chunk appears in the first k results. Its naive baseline scored 80.92% at Pass@5, 87.15% at Pass@10 and 90.06% at Pass@20. On that dataset, contextual embeddings raised Pass@10 from about 87% to about 95%. The Contextual Retrieval post measures from the other side, reporting 1 minus recall@20: the share of relevant chunks missing from the top 20. The two sets of figures come from different datasets, so compare each only with its own baseline.

    The sources also compare other ways of adding context to an index, and they don't agree on everything. The Contextual Retrieval post reports that adding generic document summaries to chunks gave very limited gains, and that summary-based indexing performed poorly in its evaluation. What worked there was context specific to each chunk. Claude's legal summarization guide takes a different view for a different situation. It says basic RAG may fall short for large documents or when precise retrieval is crucial, and it describes 'summary indexed documents': Claude writes a summary of each document, then ranks those summaries against the query. The guide calls this an advanced RAG approach that uses less context. The sources don't say whether the post's 'summary-based indexing' is the same technique. Treat them as separate options with separate evidence, and let your own Pass@k or recall numbers decide.

    Sources314

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Prepending the same document summary to every chunk is what Contextual Retrieval means.Why is that wrong?

      The context has to be specific to each chunk, placing that chunk within its document. Anthropic tried generic document summaries on chunks and saw very limited gains.

      Covered in Measuring retrieval and weighing alternative indexing ideas

    2. 2.Passing more retrieved chunks to the model always improves answers.Why is that wrong?

      More chunks raise the chance of including the relevant one, but extra material can distract the model. K is a parameter to test, and reranking is the way to get the right chunks into a small K.

      Covered in Reranking and deciding how many chunks to pass

    3. 3.Contextual Retrieval is too expensive for large corpora, because the whole document is sent once per chunk.Why is that wrong?

      With prompt caching the document is cached once and reused for each chunk. The post puts the one-time cost at about $1.02 per million document tokens.

      Covered in Giving every chunk its missing context

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “this chunk on its own doesn't specify which company it's referring to or the relevant time period”
      ↩︎ Giving every chunk its missing context
      “The resulting contextual text, usually 50-100 tokens, is prepended to the chunk before embedding it and before creating the BM25 index.”
      ↩︎ Giving every chunk its missing context
      “You simply load the document into the cache once and then reference the previously cached content.”
      ↩︎ Giving every chunk its missing context
      “including a glossary of key terms that might only be defined in other documents in the knowledge base”
      ↩︎ Giving every chunk its missing context
      “This method can reduce the number of failed retrievals by 49% and, when combined with reranking, by 67%.”
      ↩︎ Reranking and deciding how many chunks to pass
      “found using 20 to be the most performant of these options”
      ↩︎ Reranking and deciding how many chunks to pass
      “We found Gemini and Voyage embeddings to be particularly effective.”
      ↩︎ Reranking and deciding how many chunks to pass
      “We use 1 minus recall@20 as our evaluation metric”
      ↩︎ Measuring retrieval and weighing alternative indexing ideas
      “summary-based indexing (we evaluated and saw low performance)”
      ↩︎ Measuring retrieval and weighing alternative indexing ideas
      “adding generic document summaries to chunks (we experimented and saw very limited gains)”
      ↩︎ Exam trap 1
      “more information can be distracting for models so there's a limit to this”
      ↩︎ Exam trap 2
      “the one-time cost to generate contextualized chunks is $1.02 per million document tokens”
      ↩︎ Exam trap 3
    2. 2.
      “produce chunk-level vectors that capture full document context without manual metadata augmentation”
      ↩︎ Giving every chunk its missing context
      “Call these models with contextualized_embed() instead of embed()”
      ↩︎ Giving every chunk its missing context
      “take a query and a list of documents and return them ranked by relevance to the query”
      ↩︎ Reranking and deciding how many chunks to pass
      “Highest accuracy. Recommended for most applications.”
      ↩︎ Reranking and deciding how many chunks to pass
    3. 3.
      “Pass@k checks whether or not the 'golden document' was present in the first k documents retrieved for each query.”
      ↩︎ Measuring retrieval and weighing alternative indexing ideas
      “Contextual Embeddings in this case helped us to improve Pass@10 performance from ~87% --> ~95%.”
      ↩︎ Measuring retrieval and weighing alternative indexing ideas
    4. 4.
      “a basic RAG approach may be insufficient”
      ↩︎ Measuring retrieval and weighing alternative indexing ideas
      “an advanced RAG approach that provides a more efficient way of ranking documents for retrieval, using less context than traditional RAG methods”
      ↩︎ Measuring retrieval and weighing alternative indexing ideas