What you will be able to do
- Decide whether a knowledge base needs a RAG pipeline at all, or can go straight into the prompt
- Describe the preprocessing stages of a RAG pipeline: chunking, embedding and storing in a vector index
- Choose an embedding model using the vendor's selection criteria, and embed documents and queries correctly
- Explain why a hybrid index combining BM25 and embeddings retrieves better than embeddings alone
Key concept
RAG pipeline — A RAG pipeline prepares a corpus ahead of time by splitting it into chunks and indexing them. When a query arrives, it pulls the most relevant chunks and puts them in the model's prompt. Every design choice in this lesson is a choice about one of those stages.
1.First decide whether you need RAG at all
Claude's documentation describes RAG as the usual way to search a collection of documents with an LLM. Anthropic's Contextual Retrieval post adds a design rule that comes before any chunking decision: the size of the knowledge base decides whether a retrieval pipeline is worth building. If the whole corpus is smaller than about 200,000 tokens, roughly 500 pages, you can put all of it in the prompt and skip retrieval. Prompt caching makes that option cheaper and faster, because the large, repeated prompt is cached between API calls. The post reports latency reduced by more than 2x and costs cut by up to 90%.
RAG is the answer once the corpus outgrows that limit. For a knowledge base that doesn't fit in the context window, you index it once and retrieve only the pieces a given query needs. The rest of this lesson assumes you have crossed that line.
Probably not. At about 120,000 tokens it is under the roughly 200,000-token limit, so the whole handbook can go in the prompt, with prompt caching keeping repeated calls cheap and fast. RAG earns its complexity once the knowledge base no longer fits in the context window.
2.Chunking the corpus and choosing an embedding model
Preprocessing has three stages. First, split the corpus into chunks, usually no more than a few hundred tokens each. Second, turn each chunk into a vector with an embedding model. Third, store those vectors in a vector database that supports search by semantic similarity. When a query arrives, it is embedded as well, and the nearest chunks are returned.
Chunking is a real design decision. The Contextual Retrieval post says chunk size, chunk boundary and chunk overlap all affect retrieval performance. The cookbook shows two simple starting points. Its codebase dataset was split with a basic character-based splitter. Its 'naive RAG' baseline splits documents by heading, so each chunk holds the content of one subheading. Neither is presented as ideal. They are baselines you measure against.
Anthropic does not offer its own embedding model. Claude's embeddings guide lists three things to weigh when you choose a provider. The first is the model's training data: larger or more domain-specific training data generally gives better in-domain embeddings. The second is inference performance, meaning lookup speed and end-to-end latency, which matters most in large production deployments. The third is customization, such as continued training on private data. The guide then features Voyage AI, and also says to assess a variety of vendors.
| Model | Context length | Stated optimization |
|---|---|---|
| voyage-4-large | 32,000 | The best general-purpose and multilingual retrieval quality |
| voyage-4-lite | 32,000 | Optimized for latency and cost |
| voyage-code-3 | 32,000 | Optimized for code retrieval |
| voyage-finance-2 | 32,000 | Optimized for finance retrieval and RAG |
| voyage-law-2 | 16,000 | Optimized for legal and long-context retrieval and RAG |
Embed the corpus and the queries differently. Voyage takes an input_type argument: document when you index chunks, query when you embed the user's question. Because Voyage vectors are normalized, a dot product gives the cosine similarity directly.
# Embed the query
query_embd = vo.embed([query], model="voyage-4", input_type="query").embeddings[0]
# Compute the similarity
# Voyage embeddings are normalized to length 1, therefore dot-product
# and cosine similarity are the same.
similarities = np.dot(doc_embds, query_embd)
retrieved_id = np.argmax(similarities)
print(documents[retrieved_id])A search team wants their lexical index to keep matching exact product codes and error strings verbatim, while also benefiting from the same document-level grounding they just added to their semantic embeddings. They plan to keep both a BM25 index and a vector index in the pipeline. What should they do to align the two indexes?
Correct answer: A — Prepend the same chunk-specific contextual text used for embeddings to each chunk before it is tokenized for the BM25 index.
- A. Correct. This is Contextual BM25: indexing the same contextualized chunk text used for embeddings lets the lexical index capture both exact matches and the semantic relationships from the prepended context, aligning it with the contextual embeddings.
- B. Incorrect. Replacing BM25 with another embedding index removes the exact-match lexical matching (product codes, error strings) that the team explicitly wants to preserve.
- C. Incorrect. Reducing term-frequency weighting suppresses the lexical precision the team relies on for verbatim matches and does nothing to align the index with document-level context.
- D. Incorrect. Indexing only titles and headings discards the chunk body text needed for exact lexical matches on codes and error strings, worsening rather than improving alignment.
3.Indexing twice: BM25 alongside embeddings
An embedding index lets you search by meaning in vector space. That strength has a matching weakness: embedding models capture semantic relationships but can miss exact matches. The post's example is a support database queried for "Error code TS-999". An embedding search may return pages about error codes in general and miss the one page that mentions TS-999.
BM25 covers that gap. It is a lexical ranking function built on TF-IDF, and it adjusts for document length and applies saturation to term frequency, so common words don't dominate the results. It works best on queries that contain unique identifiers or technical terms. So a stronger pipeline builds two indexes over the same chunks and merges them at query time:
| Step | What happens |
|---|---|
| 1 | Split the corpus into chunks of a few hundred tokens |
| 2 | Create TF-IDF encodings and semantic embeddings for each chunk |
| 3 | BM25 finds the top chunks by exact match |
| 4 | Embeddings find the top chunks by semantic similarity |
| 5 | Combine and deduplicate results from steps 3 and 4 using rank fusion |
| 6 | Add the top-K chunks to the prompt |
The post warns that even this hybrid design has a structural limit: traditional RAG systems "often destroy context". Splitting a document can separate a chunk from the facts it depends on, such as which company or which quarter it is about. Fixing that at indexing time is what Contextual Retrieval does.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Any knowledge-grounded Claude application should be built as a RAG pipeline.Why is that wrong?
For a small corpus the simplest design wins: put the whole knowledge base in the prompt, with prompt caching to control cost and latency. RAG is for knowledge bases too large for the context window.
Covered in First decide whether you need RAG at all
2.Claude pipelines should use Anthropic's own embedding model for the vector index.Why is that wrong?
Anthropic does not offer an embedding model. You choose a third-party provider, such as Voyage AI, based on domain fit, inference performance and customization.
Covered in Chunking the corpus and choosing an embedding model
3.A good enough embedding model makes a lexical index unnecessary.Why is that wrong?
Embeddings capture meaning but can miss exact strings like error codes or identifiers. BM25 finds those, which is why the post pairs the two and merges their results with rank fusion.
Covered in Indexing twice: BM25 alongside embeddings
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Searching a collection of documents with an LLM usually involves retrieval-augmented generation (RAG).”
↩︎ First decide whether you need RAG at all - 2.https://www.anthropic.com/engineering/contextual-retrievalSecondary source
“If your knowledge base is smaller than 200,000 tokens (about 500 pages of material)”
↩︎ First decide whether you need RAG at all“For larger knowledge bases that don't fit within the context window, RAG is the typical solution.”
↩︎ First decide whether you need RAG at all“into smaller chunks of text, usually no more than a few hundred tokens”
↩︎ Chunking the corpus and choosing an embedding model“The choice of chunk size, chunk boundary, and chunk overlap can affect retrieval performance”
↩︎ Chunking the corpus and choosing an embedding model“It's particularly effective for queries that include unique identifiers or technical terms.”
↩︎ Indexing twice: BM25 alongside embeddings“Combine and deduplicate results from (3) and (4) using rank fusion techniques;”
↩︎ Indexing twice: BM25 alongside embeddings“the most relevant chunks are added to the prompt sent to the generative model”
↩︎ Key concept“you can just include the entire knowledge base in the prompt that you give the model, with no need for RAG or similar methods”
↩︎ Exam trap 1“While embedding models excel at capturing semantic relationships, they can miss crucial exact matches.”
↩︎ Exam trap 3 - 3.
“Chunk documents by heading - containing only the content from each subheading”
↩︎ Chunking the corpus and choosing an embedding model - 4.
“Larger or more domain-specific data generally produces better in-domain embeddings”
↩︎ Chunking the corpus and choosing an embedding model“are used for embedding the document and query, respectively”
↩︎ Chunking the corpus and choosing an embedding model“The embeddings allow you to do semantic search / retrieval in the vector space.”
↩︎ Indexing twice: BM25 alongside embeddings“Anthropic does not offer its own embedding model.”
↩︎ Exam trap 2