CertSafari
    CCAR-P · Lessons

    Domain 3 · Lesson 17/38

    Retrieval for Structured, Multi-Hop and Context-Poor Data: SQL, Graphs, Contextual Retrieval and Reranking

    Apply retrieval strategies matched to data shape and query pattern

    7 min read
    2.38% of exam
    6 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Route questions over structured tables to text-to-SQL instead of vector search
    • Recognise multi-hop, relationship-spanning questions that need graph traversal
    • Apply contextual retrieval when chunks lose the document context that identifies them
    • Use reranking and a tuned chunk count to trade accuracy against latency and cost

    1.Structured tables and multi-hop relationships

    Embedding retrieval works in vector space: it returns chunks whose meaning is close to the query's. Two data shapes break that model. The first is structured data. A question like "which employees are in Engineering" is answered exactly by a database query, and ranking similar text passages can only approximate it. Anthropic's text-to-SQL cookbook has Claude turn the natural-language question into SQL, given the schema in the prompt, and then runs that SQL against the database.

    SQL that Claude generated for "employees in Engineering": an exact join over two tables, which no similarity search can reproducesql
    SELECT e.name
    FROM employees e
    JOIN departments d ON e.department_id = d.id
    WHERE d.name = 'Engineering';

    Retrieval still has a role here, just over a different target. On large databases the same cookbook uses RAG to handle more complex database schemas, so the part being retrieved is the relevant schema, not the answer.

    The second shape is the multi-hop question, such as "who works with people who worked on project X". No single document contains the answer, so retrieving the best-matching chunk cannot find it. The knowledge-graph cookbook extracts entities as nodes and typed relations as edges. Multi-hop reasoning then becomes graph traversal, and Claude reads the serialized subgraph. The MongoDB fraud cookbook adds graph traversal as a fourth pattern alongside vector, full-text and hybrid search. It follows account links to surface a money-flow ring, and the cookbook calls this a relationship signal kept separate from similarity ranking.

    Four retrieval patterns on one engine (MongoDB Atlas cookbook) and the stage each uses
    PatternBuilderMongoDB stage
    Vector searchbuild_vector_pipeline$vectorSearch
    Full-text searchbuild_lexical_pipeline$search
    Hybrid (reciprocal rank fusion)build_rank_fusion_pipeline$rankFusion (8.0+)
    Graph traversalbuild_graph_pipeline$graphLookup

    A research agent runs for several hours, repeatedly retrieving and discussing documents, and the conversation is approaching its context window limit while still needing to reference earlier findings. Which strategy matches this query pattern?

    Sources123

    2.Contextual retrieval when chunks lose their context

    Some corpora are the right shape for chunk retrieval, but their chunks lose meaning once they are cut out. Anthropic's example is an SEC-filings corpus. The chunk "The company's revenue grew by 3% over the previous quarter" names neither the company nor the quarter. A query about ACME Corp in Q2 2023 therefore has nothing to match against.

    Contextual Retrieval fixes this during preprocessing. For every chunk, Claude writes a short context that places the chunk in its source document, usually 50–100 tokens. That context is prepended to the chunk before it is embedded (Contextual Embeddings) and before the BM25 index is built (Contextual BM25). Prompt caching keeps the cost down, because the whole document is cached once and reused for each of its chunks. Anthropic puts the one-time cost at $1.02 per million document tokens under its stated assumptions.

    The contextualizer prompt Anthropic used to generate per-chunk contexttext
    <document>
    {{WHOLE_DOCUMENT}}
    </document>
    Here is the chunk we want to situate within the whole document
    <chunk>
    {{CHUNK_CONTENT}}
    </chunk>
    Please give a short succinct context to situate this chunk within the overall document for the purposes of improving search retrieval of the chunk. Answer only with the succinct context and nothing else.

    The measured gains are large. Across Anthropic's test domains, Contextual Embeddings cut the top-20 retrieval failure rate by 35%, from 5.7% to 3.7%. Adding Contextual BM25 took the reduction to 49%, down to 2.9%. On the cookbook's codebase dataset, Pass@10 rose from about 87% to about 95%. There is also a model-side option: Voyage's contextualized chunk embedding models (voyage-context-3, voyage-context-4) are called with contextualized_embed() and capture document context without manual augmentation.

    Sources145

    3.Reranking and how many chunks to pass

    The last lever is what happens after first-stage retrieval. A reranker takes the query and the candidate chunks and reorders them by relevance. Anthropic reports that Contextual Retrieval alone cuts failed retrievals by 49%, and by 67% with reranking added. The reranker you choose is an accuracy–latency trade-off, and Voyage's two rerankers are positioned at opposite ends of it.

    Voyage rerankers: accuracy versus latency and cost
    ModelContext lengthPositioning
    rerank-2.532,000Highest accuracy. Recommended for most applications.
    rerank-2.5-lite32,000Optimized for latency and cost.

    The number of chunks passed to Claude also needs tuning. Passing more chunks makes it more likely the relevant one is included. But extra material can distract the model, so there is a ceiling. Anthropic tested 5, 10 and 20 chunks and got the best results with 20, but advises testing on your own use case. Chunk size, chunk boundaries and overlap affect retrieval performance too, and so does the embedding model: Anthropic found Gemini and Voyage embeddings particularly effective.

    A platform wants Claude to discover tools using an existing embeddings index rather than the built-in regex or BM25 matching, because tool descriptions are sparse but semantically related through the embedding space. Which implementation matches this requirement?

    Sources14

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Standard chunk-based RAG can answer questions whose answer is spread across several documents, if you retrieve enough chunks.Why is that wrong?

      Chunk retrieval returns similar passages. It doesn't chain facts from one passage to the next. For multi-hop, relationship questions, build a knowledge graph and traverse it.

      Covered in Structured tables and multi-hop relationships

    2. 2.Adding a generic document summary to each chunk is the same technique as Contextual Retrieval.Why is that wrong?

      Contextual Retrieval prepends context written for each individual chunk. Anthropic tested generic summaries and saw very limited gains.

      Covered in Contextual retrieval when chunks lose their context

    3. 3.Passing more retrieved chunks to the model always improves answers.Why is that wrong?

      More chunks raise recall, but extra material can distract the model, so there is a limit. Tune the chunk count on your own use case.

      Covered in Reranking and how many chunks to pass

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The embeddings allow you to do semantic search / retrieval in the vector space.”
      ↩︎ Structured tables and multi-hop relationships
      “The following contextualized chunk embedding models produce chunk-level vectors that capture full document context without manual metadata augmentation.”
      ↩︎ Contextual retrieval when chunks lose their context
      “Voyage AI also offers rerankers, which take a query and a list of documents and return them ranked by relevance to the query.”
      ↩︎ Reranking and how many chunks to pass
    2. 2.
      “Text to SQL is a natural language processing task that converts human-readable text queries into structured SQL queries.”
      ↩︎ Structured tables and multi-hop relationships
      “RAG (Retrieval Augmented Generation) to handle more complex database systems”
      ↩︎ Structured tables and multi-hop relationships
    3. 3.
      “This is a relationship signal, deliberately separate from the similarity ranking above.”
      ↩︎ Structured tables and multi-hop relationships
    4. 4.
      “Contextual Retrieval solves this problem by prepending chunk-specific explanatory context to each chunk before embedding”
      ↩︎ Contextual retrieval when chunks lose their context
      “The resulting contextual text, usually 50-100 tokens, is prepended to the chunk before embedding it and before creating the BM25 index.”
      ↩︎ Contextual retrieval when chunks lose their context
      “the one-time cost to generate contextualized chunks is $1.02 per million document tokens”
      ↩︎ Contextual retrieval when chunks lose their context
      “This method can reduce the number of failed retrievals by 49% and, when combined with reranking, by 67%.”
      ↩︎ Reranking and how many chunks to pass
      “We tried delivering 5, 10, and 20 chunks, and found using 20 to be the most performant of these options”
      ↩︎ Reranking and how many chunks to pass
      “adding generic document summaries to chunks (we experimented and saw very limited gains)”
      ↩︎ Exam trap 2
      “However, more information can be distracting for models”
      ↩︎ Exam trap 3
    5. 5.
      “Contextual Embeddings in this case helped us to improve Pass@10 performance from ~87% --> ~95%.”
      ↩︎ Contextual retrieval when chunks lose their context

    Also cited