What you will be able to do
- Explain why isolated chunks retrieve poorly, and how contextual embeddings and contextual BM25 fix that at indexing time
- Estimate the cost of contextualizing a corpus, and explain the role prompt caching plays
- Add a reranking stage and choose how many chunks to pass to the model
- Evaluate retrieval with Pass@k or recall@k, and weigh the alternative indexing ideas the sources compare
1.Giving every chunk its missing context
Standard RAG splits documents into chunks of a few hundred tokens before indexing them. That works until a chunk depends on context it no longer carries. Anthropic's example is a set of SEC filings queried with "What was the revenue growth for ACME Corp in Q2 2023?" The right chunk reads: "The company's revenue grew by 3% over the previous quarter." On its own, that chunk doesn't say which company or which period it describes. Neither the embedding index nor the keyword index has anything linking it to ACME or Q2 2023.
Contextual Retrieval fixes this during preprocessing. It prepends a short, chunk-specific explanation to each chunk before the chunk is embedded (Contextual Embeddings) and before it goes into the BM25 index (Contextual BM25). Writing that context by hand for millions of chunks isn't practical, so Claude writes it. The post used this prompt with Claude 3 Haiku:
<document>
{{WHOLE_DOCUMENT}}
</document>
Here is the chunk we want to situate within the whole document
<chunk>
{{CHUNK_CONTENT}}
</chunk>
Please give a short succinct context to situate this chunk within the overall document for the purposes of improving search retrieval of the chunk. Answer only with the succinct context and nothing else.The generated context is usually 50 to 100 tokens. Sending the whole document once for every chunk sounds expensive, and prompt caching is what makes it affordable: the document is cached once and each chunk's request reads from that cache. With 800-token chunks and 8k-token documents, the post puts the one-time cost at $1.02 per million document tokens. The generic prompt works well, but a domain-specific one can do better, for example one that includes a glossary of terms defined elsewhere in the corpus.
The embedding layer offers a second route. Claude's embeddings guide lists Voyage's contextualized chunk embedding models, voyage-context-4 and voyage-context-3, with a 120,000-token context length. You call them with contextualized_embed() instead of embed(). They produce chunk vectors that already reflect the full document, with no manual metadata augmentation.
A support-bot pipeline retrieves the same set of product documentation chunks for every user session and appends them to the system prompt, followed by whatever question the current user asks. The team wants to use prompt caching to cut latency and cost across the many queries that reuse this documentation. Where should they place the cache breakpoint?
Correct answer: A — On the last content block of the retrieved documentation, immediately before the user's question is added to the request.
- A. Correct. The cache breakpoint should sit on the last block that stays identical across requests. In this pipeline, that is the end of the retrieved documentation, right before the variable user question, so every query benefits from a cache hit on the stable content.
- B. Incorrect. The user's question changes on every request, so placing the breakpoint there means the cached prefix never matches between requests and no cache hits occur.
- C. Incorrect. Caching only the first sentence of the system prompt leaves the much larger, genuinely stable retrieved documentation uncached, missing almost all of the available savings.
- D. Incorrect. Caching does not require the entire request to be identical; only the cached prefix (here, the retrieved documentation) needs to match, so this pipeline is a strong caching candidate.
2.Reranking and deciding how many chunks to pass
The indexing changes improve recall. A reranking stage improves the order of what comes back. A reranker takes the query and the candidate chunks and returns them sorted by relevance, so the chunks most worth including rise to the top before you cut the list down. Claude's embeddings guide lists two Voyage rerankers, both called with rerank(): rerank-2.5 for highest accuracy, recommended for most applications, and rerank-2.5-lite, optimized for latency and cost.
| Configuration | Failure rate | Reduction vs. baseline |
|---|---|---|
| Standard embeddings (baseline) | 5.7% | — |
| Contextual Embeddings | 3.7% | 35% |
| Contextual Embeddings + Contextual BM25 | 2.9% | 49% |
| Contextual Embeddings + Contextual BM25 + reranking | Not stated as an absolute rate | 67% |
The last knob is K, the number of chunks you put in the prompt. More chunks raise the odds that the relevant one is included, but the post warns that extra material can distract the model. Of the 5, 10 and 20 chunks it tested, 20 performed best, and it recommends experimenting on your own use case rather than treating 20 as a rule. The post also found that contextualizing helped with every embedding model it tested, with Gemini and Voyage embeddings especially effective.
Add a reranker between retrieval and the prompt. It re-sorts the candidate set by relevance to the query, so the golden chunk moves into the top 20. Raising K to 50 adds material that can distract the model. The post found diminishing returns and recommends testing K rather than assuming more is better.
A platform team is redesigning their RAG pipeline after an audit found a high rate of retrieval failures, where relevant chunks exist in the knowledge base but are not returned for matching queries. They can implement several changes before the next release. Which of the following changes are documented to measurably reduce retrieval failure rates? (Select all that apply.)(Select 3)
Correct answers: A, B, C — Prepend chunk-specific contextual summaries to chunks before generating their embeddings.; Index the same contextualized chunk text lexically with BM25 alongside the embeddings.; Add a reranking stage that rescores initial retrieval candidates before generation.
- A. Correct. Contextual embeddings, which prepend explanatory context before encoding, are documented to reduce retrieval failures on their own.
- B. Correct. Combining contextual embeddings with contextual BM25 lexical indexing further reduces retrieval failures beyond embeddings alone.
- C. Correct. Adding a reranking stage on top of contextual embeddings and BM25 produces the largest documented reduction in retrieval failures of the three combined techniques.
- D. Incorrect. Sampling temperature affects the wording and variability of the generated answer text; it has no effect on which chunks are retrieved or how retrieval failures occur.
- E. Incorrect. Binary quantization is a storage and cost optimization that reduces embedding precision; it is not a documented technique for reducing retrieval failure rates and can trade off retrieval accuracy for space.
- F. Incorrect. Disabling overlap speeds up index building but removes a safeguard against split-boundary content, which tends to increase rather than decrease retrieval failures.
3.Measuring retrieval and weighing alternative indexing ideas
Every chunking, indexing or reranking change should be measured against a baseline. The contextual embeddings cookbook builds an evaluation set of 248 queries, each paired with a 'golden chunk', and scores Pass@k: whether that chunk appears in the first k results. Its naive baseline scored 80.92% at Pass@5, 87.15% at Pass@10 and 90.06% at Pass@20. On that dataset, contextual embeddings raised Pass@10 from about 87% to about 95%. The Contextual Retrieval post measures from the other side, reporting 1 minus recall@20: the share of relevant chunks missing from the top 20. The two sets of figures come from different datasets, so compare each only with its own baseline.
The sources also compare other ways of adding context to an index, and they don't agree on everything. The Contextual Retrieval post reports that adding generic document summaries to chunks gave very limited gains, and that summary-based indexing performed poorly in its evaluation. What worked there was context specific to each chunk. Claude's legal summarization guide takes a different view for a different situation. It says basic RAG may fall short for large documents or when precise retrieval is crucial, and it describes 'summary indexed documents': Claude writes a summary of each document, then ranks those summaries against the query. The guide calls this an advanced RAG approach that uses less context. The sources don't say whether the post's 'summary-based indexing' is the same technique. Treat them as separate options with separate evidence, and let your own Pass@k or recall numbers decide.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Prepending the same document summary to every chunk is what Contextual Retrieval means.Why is that wrong?
The context has to be specific to each chunk, placing that chunk within its document. Anthropic tried generic document summaries on chunks and saw very limited gains.
Covered in Measuring retrieval and weighing alternative indexing ideas
2.Passing more retrieved chunks to the model always improves answers.Why is that wrong?
More chunks raise the chance of including the relevant one, but extra material can distract the model. K is a parameter to test, and reranking is the way to get the right chunks into a small K.
3.Contextual Retrieval is too expensive for large corpora, because the whole document is sent once per chunk.Why is that wrong?
With prompt caching the document is cached once and reused for each chunk. The post puts the one-time cost at about $1.02 per million document tokens.
Covered in Giving every chunk its missing context
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://www.anthropic.com/engineering/contextual-retrievalSecondary source
“this chunk on its own doesn't specify which company it's referring to or the relevant time period”
↩︎ Giving every chunk its missing context“The resulting contextual text, usually 50-100 tokens, is prepended to the chunk before embedding it and before creating the BM25 index.”
↩︎ Giving every chunk its missing context“You simply load the document into the cache once and then reference the previously cached content.”
↩︎ Giving every chunk its missing context“including a glossary of key terms that might only be defined in other documents in the knowledge base”
↩︎ Giving every chunk its missing context“This method can reduce the number of failed retrievals by 49% and, when combined with reranking, by 67%.”
↩︎ Reranking and deciding how many chunks to pass“found using 20 to be the most performant of these options”
↩︎ Reranking and deciding how many chunks to pass“We found Gemini and Voyage embeddings to be particularly effective.”
↩︎ Reranking and deciding how many chunks to pass“We use 1 minus recall@20 as our evaluation metric”
↩︎ Measuring retrieval and weighing alternative indexing ideas“summary-based indexing (we evaluated and saw low performance)”
↩︎ Measuring retrieval and weighing alternative indexing ideas“adding generic document summaries to chunks (we experimented and saw very limited gains)”
↩︎ Exam trap 1“more information can be distracting for models so there's a limit to this”
↩︎ Exam trap 2“the one-time cost to generate contextualized chunks is $1.02 per million document tokens”
↩︎ Exam trap 3 - 2.
“produce chunk-level vectors that capture full document context without manual metadata augmentation”
↩︎ Giving every chunk its missing context“Call these models with contextualized_embed() instead of embed()”
↩︎ Giving every chunk its missing context“take a query and a list of documents and return them ranked by relevance to the query”
↩︎ Reranking and deciding how many chunks to pass“Highest accuracy. Recommended for most applications.”
↩︎ Reranking and deciding how many chunks to pass - 3.
“Pass@k checks whether or not the 'golden document' was present in the first k documents retrieved for each query.”
↩︎ Measuring retrieval and weighing alternative indexing ideas“Contextual Embeddings in this case helped us to improve Pass@10 performance from ~87% --> ~95%.”
↩︎ Measuring retrieval and weighing alternative indexing ideas - 4.
“a basic RAG approach may be insufficient”
↩︎ Measuring retrieval and weighing alternative indexing ideas“an advanced RAG approach that provides a more efficient way of ranking documents for retrieval, using less context than traditional RAG methods”
↩︎ Measuring retrieval and weighing alternative indexing ideas