CertSafari
    CCAR-P · Lessons

    Domain 3 · Lesson 16/38

    RAG Pipeline Design: Chunking, Embedding and Hybrid Indexing

    Design a RAG pipeline with appropriate chunking and indexing strategies

    7 min read
    2.38% of exam
    4 sources
    Published 27 Sep 2026
    Docs as of 26 Sep 2026

    What you will be able to do

    • Decide whether a knowledge base needs a RAG pipeline at all, or can go straight into the prompt
    • Describe the preprocessing stages of a RAG pipeline: chunking, embedding and storing in a vector index
    • Choose an embedding model using the vendor's selection criteria, and embed documents and queries correctly
    • Explain why a hybrid index combining BM25 and embeddings retrieves better than embeddings alone

    Key concept

    RAG pipeline — A RAG pipeline prepares a corpus ahead of time by splitting it into chunks and indexing them. When a query arrives, it pulls the most relevant chunks and puts them in the model's prompt. Every design choice in this lesson is a choice about one of those stages.

    1.First decide whether you need RAG at all

    Claude's documentation describes RAG as the usual way to search a collection of documents with an LLM. Anthropic's Contextual Retrieval post adds a design rule that comes before any chunking decision: the size of the knowledge base decides whether a retrieval pipeline is worth building. If the whole corpus is smaller than about 200,000 tokens, roughly 500 pages, you can put all of it in the prompt and skip retrieval. Prompt caching makes that option cheaper and faster, because the large, repeated prompt is cached between API calls. The post reports latency reduced by more than 2x and costs cut by up to 90%.

    RAG is the answer once the corpus outgrows that limit. For a knowledge base that doesn't fit in the context window, you index it once and retrieve only the pieces a given query needs. The rest of this lesson assumes you have crossed that line.

    Sources12

    2.Chunking the corpus and choosing an embedding model

    Preprocessing has three stages. First, split the corpus into chunks, usually no more than a few hundred tokens each. Second, turn each chunk into a vector with an embedding model. Third, store those vectors in a vector database that supports search by semantic similarity. When a query arrives, it is embedded as well, and the nearest chunks are returned.

    Chunking is a real design decision. The Contextual Retrieval post says chunk size, chunk boundary and chunk overlap all affect retrieval performance. The cookbook shows two simple starting points. Its codebase dataset was split with a basic character-based splitter. Its 'naive RAG' baseline splits documents by heading, so each chunk holds the content of one subheading. Neither is presented as ideal. They are baselines you measure against.

    Anthropic does not offer its own embedding model. Claude's embeddings guide lists three things to weigh when you choose a provider. The first is the model's training data: larger or more domain-specific training data generally gives better in-domain embeddings. The second is inference performance, meaning lookup speed and end-to-end latency, which matters most in large production deployments. The third is customization, such as continued training on private data. The guide then features Voyage AI, and also says to assess a variety of vendors.

    Choosing a Voyage embedding model for the kind of content you are indexing
    ModelContext lengthStated optimization
    voyage-4-large32,000The best general-purpose and multilingual retrieval quality
    voyage-4-lite32,000Optimized for latency and cost
    voyage-code-332,000Optimized for code retrieval
    voyage-finance-232,000Optimized for finance retrieval and RAG
    voyage-law-216,000Optimized for legal and long-context retrieval and RAG

    Embed the corpus and the queries differently. Voyage takes an input_type argument: document when you index chunks, query when you embed the user's question. Because Voyage vectors are normalized, a dot product gives the cosine similarity directly.

    Embedding a query with input_type="query" and ranking the pre-embedded documents by similaritypython
    # Embed the query
    query_embd = vo.embed([query], model="voyage-4", input_type="query").embeddings[0]
    
    # Compute the similarity
    # Voyage embeddings are normalized to length 1, therefore dot-product
    # and cosine similarity are the same.
    similarities = np.dot(doc_embds, query_embd)
    
    retrieved_id = np.argmax(similarities)
    print(documents[retrieved_id])

    A search team wants their lexical index to keep matching exact product codes and error strings verbatim, while also benefiting from the same document-level grounding they just added to their semantic embeddings. They plan to keep both a BM25 index and a vector index in the pipeline. What should they do to align the two indexes?

    Sources234

    3.Indexing twice: BM25 alongside embeddings

    An embedding index lets you search by meaning in vector space. That strength has a matching weakness: embedding models capture semantic relationships but can miss exact matches. The post's example is a support database queried for "Error code TS-999". An embedding search may return pages about error codes in general and miss the one page that mentions TS-999.

    BM25 covers that gap. It is a lexical ranking function built on TF-IDF, and it adjusts for document length and applies saturation to term frequency, so common words don't dominate the results. It works best on queries that contain unique identifiers or technical terms. So a stronger pipeline builds two indexes over the same chunks and merges them at query time:

    The hybrid retrieval flow described in the Contextual Retrieval post
    StepWhat happens
    1Split the corpus into chunks of a few hundred tokens
    2Create TF-IDF encodings and semantic embeddings for each chunk
    3BM25 finds the top chunks by exact match
    4Embeddings find the top chunks by semantic similarity
    5Combine and deduplicate results from steps 3 and 4 using rank fusion
    6Add the top-K chunks to the prompt

    The post warns that even this hybrid design has a structural limit: traditional RAG systems "often destroy context". Splitting a document can separate a chunk from the facts it depends on, such as which company or which quarter it is about. Fixing that at indexing time is what Contextual Retrieval does.

    Sources42

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Any knowledge-grounded Claude application should be built as a RAG pipeline.Why is that wrong?

      For a small corpus the simplest design wins: put the whole knowledge base in the prompt, with prompt caching to control cost and latency. RAG is for knowledge bases too large for the context window.

      Covered in First decide whether you need RAG at all

    2. 2.Claude pipelines should use Anthropic's own embedding model for the vector index.Why is that wrong?

      Anthropic does not offer an embedding model. You choose a third-party provider, such as Voyage AI, based on domain fit, inference performance and customization.

      Covered in Chunking the corpus and choosing an embedding model

    3. 3.A good enough embedding model makes a lexical index unnecessary.Why is that wrong?

      Embeddings capture meaning but can miss exact strings like error codes or identifiers. BM25 finds those, which is why the post pairs the two and merges their results with rank fusion.

      Covered in Indexing twice: BM25 alongside embeddings

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Searching a collection of documents with an LLM usually involves retrieval-augmented generation (RAG).”
      ↩︎ First decide whether you need RAG at all
    2. 2.
      “If your knowledge base is smaller than 200,000 tokens (about 500 pages of material)”
      ↩︎ First decide whether you need RAG at all
      “For larger knowledge bases that don't fit within the context window, RAG is the typical solution.”
      ↩︎ First decide whether you need RAG at all
      “into smaller chunks of text, usually no more than a few hundred tokens”
      ↩︎ Chunking the corpus and choosing an embedding model
      “The choice of chunk size, chunk boundary, and chunk overlap can affect retrieval performance”
      ↩︎ Chunking the corpus and choosing an embedding model
      “It's particularly effective for queries that include unique identifiers or technical terms.”
      ↩︎ Indexing twice: BM25 alongside embeddings
      “Combine and deduplicate results from (3) and (4) using rank fusion techniques;”
      ↩︎ Indexing twice: BM25 alongside embeddings
      “the most relevant chunks are added to the prompt sent to the generative model”
      ↩︎ Key concept
      “you can just include the entire knowledge base in the prompt that you give the model, with no need for RAG or similar methods”
      ↩︎ Exam trap 1
      “While embedding models excel at capturing semantic relationships, they can miss crucial exact matches.”
      ↩︎ Exam trap 3
    3. 3.
      “Chunk documents by heading - containing only the content from each subheading”
      ↩︎ Chunking the corpus and choosing an embedding model
    4. 4.
      “Larger or more domain-specific data generally produces better in-domain embeddings”
      ↩︎ Chunking the corpus and choosing an embedding model
      “are used for embedding the document and query, respectively”
      ↩︎ Chunking the corpus and choosing an embedding model
      “The embeddings allow you to do semantic search / retrieval in the vector space.”
      ↩︎ Indexing twice: BM25 alongside embeddings
      “Anthropic does not offer its own embedding model.”
      ↩︎ Exam trap 2

    Continue to page 2 of 2

    Contextual Retrieval and Reranking for RAG Indexes