What you will be able to do
- Decide when a knowledge base is small enough to put straight into the prompt instead of building a retrieval pipeline
- Explain what embedding-based semantic search does well and where it fails
- Pick lexical (BM25) or hybrid retrieval when queries contain exact identifiers, codes or technical terms
- Describe how rank fusion combines semantic and lexical results into one ranked list
Key concept
Matching retrieval to the query — No single retrieval method suits every query. Embeddings find passages with similar meaning, and lexical matching finds exact strings. How big the corpus is, and what a typical query looks like, decide which method you need, whether to combine them, or whether you need retrieval at all.
1.First question: do you need retrieval at all?
Before you choose between retrieval strategies, check how big the corpus is. Anthropic's guidance puts the line at about 200,000 tokens, roughly 500 pages. Below that, you can put the entire knowledge base in the prompt and skip RAG. Nothing gets lost to a failed search, because nothing is searched. Prompt caching is what makes this affordable on repeated calls: a cached prompt is reused between API calls, which Anthropic reports cuts latency by more than 2x and costs by up to 90%.
Retrieval becomes necessary once the corpus no longer fits the context window. Claude Projects shows the same trade-off as a product. A project runs on in-context processing by default. It switches to a RAG mode automatically when its knowledge approaches the context window limit. In that mode Claude calls a project knowledge search tool instead of loading every file. If the knowledge later shrinks below the threshold, the project switches back to in-context processing.
About 300 pages is under the roughly 500-page (200,000-token) line, so the simplest correct design is to put the whole handbook in the prompt and cache it. Retrieval adds a failure mode, a relevant chunk that never gets retrieved, which the in-context approach doesn't have. Move to RAG only when the corpus outgrows the context window.
2.Semantic search: retrieval by meaning
Standard RAG preprocesses the corpus in three steps. It splits the documents into chunks, usually no more than a few hundred tokens each. It converts each chunk into a vector embedding that encodes its meaning. It stores those vectors in a database that supports similarity search. At query time the query is embedded the same way, and the chunks whose vectors lie nearest to it are added to the prompt.
The Claude embeddings documentation shows this with Voyage. Documents are embedded with input_type="document" and the query with input_type="query". The query "When is Apple's conference call scheduled?" returns the one document about Apple's earnings call, even though the wording differs. Semantic search handles exactly this kind of query: the user describes a concept and the relevant text says it in other words.
# Embed the query
query_embd = vo.embed([query], model="voyage-4", input_type="query").embeddings[0]
# Compute the similarity
# Voyage embeddings are normalized to length 1, therefore dot-product
# and cosine similarity are the same.
similarities = np.dot(doc_embds, query_embd)
retrieved_id = np.argmax(similarities)
print(documents[retrieved_id])Embeddings have a known blind spot. They are good at capturing semantic relationships, but they can miss exact matches. The contextual retrieval post gives an example: a query for "Error code TS-999" against a support database. An embedding model may return content about error codes in general and still miss the one page that mentions TS-999.
3.Lexical and hybrid retrieval for identifiers and exact terms
BM25 is a ranking function built on TF-IDF. It scores chunks by lexical matches on exact words and phrases. It also adjusts for document length and saturates term frequency, so common words don't dominate the results. In the TS-999 case, BM25 searches for that exact string and finds the right page. The MongoDB fraud-review cookbook describes full-text search the same way: it catches the exact names, IDs and codes that embeddings blur.
Real query traffic mixes both kinds of query, so the standard design runs both retrievers and merges them. You build TF-IDF encodings and embeddings for every chunk. BM25 takes the top exact matches, embeddings take the top semantic matches, and rank fusion combines and deduplicates the two lists. The top-K chunks then go into the prompt. In the MongoDB example, $rankFusion does the merge on the server by reciprocal rank, and you can weight each input to favour semantic or lexical results.
| Query pattern | Example | Retriever that fits | Why |
|---|---|---|---|
| Concept described in the user's own words | "When is Apple's conference call scheduled?" | Embeddings (semantic similarity) | Matches meaning even when the wording differs |
| Unique identifier, code or technical term | "Error code TS-999" | BM25 (lexical matching) | Finds the exact string, which embeddings can miss |
| A mix of both across real traffic | Support questions with and without IDs | Hybrid: BM25 + embeddings, merged by rank fusion | Balances precise term matching with broader semantic understanding |
An internal ops agent has 8 tools total (deploy, rollback, restart, status, logs, scale, alert, notify), each under 100 tokens, and nearly every request needs several of them. Which tool-configuration approach fits this scenario?
Correct answer: A — Call all 8 tools directly through standard tool definitions without enabling tool search or deferred loading.
- A. Correct. With fewer than 10 tools, small definitions, and near-universal use per request, standard tool calling without tool search is the better fit; deferring these tools would add search overhead for no benefit.
- B. Tool search helps when catalogs are large or selection accuracy is degrading; forcing a search step for 8 tools used in nearly every request adds latency without solving a real problem.
- C. Deferring all 8 tools behind regex search removes the 3-5 frequently used tools from immediate context, which is discouraged when most of the catalog is needed on almost every call.
- D. Deferring the whole toolset behind one MCP connector still forces a search step before nearly every request can proceed, which is unnecessary at this small scale.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Any knowledge base worth querying needs a RAG pipeline with chunking and a vector database.Why is that wrong?
A corpus under about 200,000 tokens can go into the prompt whole, with prompt caching keeping repeated calls fast and cheap. RAG is for corpora that don't fit.
2.A good embedding model makes keyword search redundant, even for queries full of error codes and IDs.Why is that wrong?
Embeddings can blur exact identifiers. BM25 matches the literal string, so queries heavy in identifiers need lexical or hybrid retrieval.
Covered in Lexical and hybrid retrieval for identifiers and exact terms
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://support.claude.com/en/articles/11473015-retrieval-augmented-generation-rag-for-projectsOfficial docs
“When possible, projects will use in-context processing for optimal performance.”
↩︎ First question: do you need retrieval at all?“When your project knowledge approaches the context window limit, Claude will automatically enable RAG mode”
↩︎ First question: do you need retrieval at all?“If your project knowledge later drops below the context window threshold, Claude can automatically convert back to context-based processing.”
↩︎ First question: do you need retrieval at all? - 2.https://www.anthropic.com/engineering/contextual-retrievalSecondary source
“Developers can now cache frequently used prompts between API calls, reducing latency by > 2x and costs by up to 90%”
↩︎ First question: do you need retrieval at all?“While embedding models excel at capturing semantic relationships, they can miss crucial exact matches.”
↩︎ Semantic search: retrieval by meaning“BM25 (Best Matching 25) is a ranking function that uses lexical matching to find precise word or phrase matches.”
↩︎ Lexical and hybrid retrieval for identifiers and exact terms“Combine and deduplicate results from (3) and (4) using rank fusion techniques”
↩︎ Lexical and hybrid retrieval for identifiers and exact terms“particularly effective for queries that include unique identifiers or technical terms”
↩︎ Key concept“If your knowledge base is smaller than 200,000 tokens (about 500 pages of material), you can just include the entire knowledge base”
↩︎ Exam trap 1“BM25 looks for this specific text string to identify the relevant documentation.”
↩︎ Exam trap 2 - 3.
“The embeddings allow you to do semantic search / retrieval in the vector space.”
↩︎ Semantic search: retrieval by meaning“conduct a nearest neighbor search to find the most relevant document based on the distance in the embedding space”
↩︎ Lexical and hybrid retrieval for identifiers and exact terms - 4.
“this is what catches the exact names, IDs, and codes that embeddings blur”
↩︎ Lexical and hybrid retrieval for identifiers and exact terms