What you will be able to do
- Explain why an evaluation set must exist before you compare chunking options
- Interpret chunk precision, document recall and context sufficiency, and know which need ground truth
- Choose which chunk-size direction to test first based on whether the use case needs recall or precision
- Tell retrieval failures that chunking can fix apart from failures caused elsewhere in the RAG chain
1.Measure before you re-chunk
You can't choose a chunking strategy well by reading chunks by eye. The AI Search retrieval quality guide says that before any optimization you need a reproducible evaluation system, and if you don't have one, you should build it first. It doesn't have to be elaborate. Any of these works: an existing golden dataset of labeled query-answer pairs, a synthetic evaluation set generated from your documents, or ground-truth-free evaluation using Agent Evaluation judges. The point is to have an automated way to measure what changes when you swap one chunking configuration for another.
The guide also asks you to set latency targets before you start: a time-to-first-token target for RAG agents, and an end-to-end display target for search bars. Any chunking or retrieval change you try has to meet them. When results come from the AI Search retrieval quality evaluation, every metric includes a 95% confidence interval. That lets you check whether the gap between two chunking runs is statistically meaningful or just noise.
Checkpoint 1 of 5· Check yourself
A team wants to try semantic chunking but has no evaluation in place. What does the Databricks guide recommend?
Without measurement you can't tell whether a chunking change helped, so the guide says to stop and build evaluation before optimizing.
“If you don't have evaluation in place, stop here and set it up first. Optimizing without measurement is guesswork.”Source: docs.databricks.com
2.The retrieval metrics that judge chunking
Chunking changes what the retriever returns, so retrieval metrics are the direct way to measure it. Agent Evaluation provides three retrieval metrics. They differ in how they are computed and in whether they need ground truth.
| Metric | Question it answers | Measured by | Needs ground truth? |
|---|---|---|---|
| chunk_relevance/precision | What % of the retrieved chunks are relevant to the request? | LLM judge | No |
| document_recall | What % of the ground truth documents are represented in the retrieved chunks? | Deterministic | Yes |
| context_sufficiency | Are the retrieved chunks sufficiency to produce the expected response? | LLM Judge | Yes |
The cookbook gives a worked example. Three chunks are retrieved and two of them are relevant, so precision is 2/3, about 0.66. Four documents are relevant in total and the retrieved chunks come from two of them, so recall is 2/4 = 0.5. You can compute precision without knowing every relevant item. Recall needs ground truth that lists all of them. That's why document_recall needs a labeled evaluation set and chunk precision doesn't.
Checkpoint 2 of 5· Check yourself
Your evaluation set has queries but no labeled relevant documents. Which retrieval metric can you still use to compare two chunking strategies?
An LLM judge scores chunk precision by judging each retrieved chunk against the request, so it needs no labels. Recall and context sufficiency both need ground truth.
“Computing precision does not require knowing all relevant items.”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
A team is preparing chunks for an embedding model whose maximum input context length is 512 tokens; any input beyond that length is truncated before it is embedded. Their current chunking pipeline splits only on double newlines between long, unstructured paragraphs, producing chunks that average 1,800 tokens. Which change should the team make to the chunking strategy so the embedding step reflects the full content of each chunk?
Correct answer: B — Reduce the target chunk size so it reliably fits within the 512-token embedding limit, using recursive splitting on long paragraphs.
- A. Changing the split unit to sentences while keeping the 1,800-token target does nothing to shrink chunks below the 512-token embedding limit, so truncation still occurs.
- B. Reducing the target chunk size so it fits under 512 tokens, using recursive splitting to break long paragraphs into smaller pieces, ensures the full chunk content is actually embedded instead of silently truncated.
- C. Batch size controls how many chunks are sent per API call, not how many tokens of a single chunk the model can accept, so this has no effect on truncation of individual 1,800-token chunks.
- D. Re-ranking operates on chunks after they have already been embedded, so it cannot recover information that was cut off before the embedding step even ran.
Sources3
3.Turning results into a chunking choice
Which metric to prioritise depends on the use case. For RAG agents, recall usually matters most, because missing key context leads to incorrect answers or hallucinations. Track Recall@k. Precision matters most when users need the exact answer at the top of the results, for example in entity matching or a tech-support knowledge base. Track Precision@k. Some use cases, such as M&A due diligence, need both.
Then link the priority to the chunk-size trade-off. The guide gives three starting configurations, and each comment states its strength. Smaller chunks are better for precise fact retrieval. Larger chunks keep more context per chunk. Run each configuration against the same evaluation set, compare the metric your use case cares about, and keep the configuration that improves it within your latency targets.
# Common configurations to test
small_chunks = 256 # Better for precise fact retrieval
medium_chunks = 512 # Balanced approach
large_chunks = 1024 # More context per chunkThe guide says larger chunks keep more context but make the relevant information harder to pinpoint, while smaller chunks localise specific information better. For a use case that needs the exact solution in the top results, the smaller configuration is the one to keep, provided it meets your latency targets.
Checkpoint 4 of 5· Exam question
A team runs a retrieval quality evaluation across three chunking strategies applied to the same product-manual corpus: 200-token fixed-size chunks, 400-token fixed-size chunks, and semantic chunking that groups sentences by topic. Using the same generated test queries and LLM-judge relevance scoring for all three, the evaluation reports DCG@10 scores of 4.1, 5.0, and 7.3 respectively. Based on this evaluation, which chunking strategy should the team select for production, and why?
Correct answer: C — The semantic chunking strategy, because it produced the highest DCG@10 score, indicating both more relevant results and better ranking of those results among the top 10.
- A. There is no rule that a smaller score gap between two options makes one automatically preferred over a third option with a clearly higher score; the 400-token result is simply lower than the semantic result.
- B. Choosing based on cost while ignoring the reported DCG@10 scores contradicts the premise of selecting a strategy based on retrieval evaluation results, and the 200-token strategy scored the lowest of the three.
- C. DCG@10 weights both whether results are relevant and where they rank in the top 10, so the semantic strategy's score of 7.3 versus 4.1 and 5.0 shows it retrieved more relevant chunks in better positions.
- D. DCG@10 explicitly discounts gain by rank position, so it does account for ranking, and a difference this large (7.3 versus 4.1 and 5.0) reflects a meaningful gap in retrieval quality.
4.Is chunking really the cause?
A low retrieval score doesn't automatically mean you should re-chunk. The data pipeline affects retrieval quality through parsing, chunking, metadata and the embedding model. The RAG chain also affects it, through query transformation, the number of chunks retrieved and re-ranking. Before you re-chunk, look at what the failing queries have in common. If the relevant text was never captured in a chunk, chunking is the likely cause. If the right chunks exist in the index but rank too low, the chain settings are the likely cause.
Response metrics and retrieval metrics also need to be read together. A RAG application can retrieve the correct context and still answer badly, and it can give a good answer from poor retrieval. Measuring both, and evaluating each component as well as the whole system, tells you whether a chunking change caused an improvement or just happened alongside one.
Checkpoint 5 of 5· Check yourself
After re-chunking, document_recall goes up but answer correctness stays flat. What is the most defensible conclusion?
Retrieval and response quality can move independently. Better recall shows the retriever improved, so look for the remaining gap on the generation side.
“A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Document recall can be computed by an LLM judge without labeled data.Why is that wrong?
document_recall is computed deterministically and needs ground truth that lists every relevant document. Only chunk precision works without labels.
Covered in The retrieval metrics that judge chunking
2.Poor retrieval metrics always mean the chunking strategy must change.Why is that wrong?
Retrieval quality also depends on chain settings such as query transformation, the number of chunks retrieved and re-ranking, so rule those out before re-chunking.
Covered in Is chunking really the cause?
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“If you don't have evaluation in place, stop here and set it up first. Optimizing without measurement is guesswork.”
↩︎ Measure before you re-chunk“RAG agents: Missing key context leads to incorrect answers or hallucinations.”
↩︎ Turning results into a chunking choice“Tech support knowledge base: Engineers need the exact solution in top results.”
↩︎ Turning results into a chunking choice“Focus on relative improvements as you test different strategies, not absolute scores.”
↩︎ Prediction - 2.
“All metrics include 95% confidence intervals computed across per-query values”
↩︎ Measure before you re-chunk“When completeness is important, such as in RAG applications where missing a relevant document means the LLM generates an incomplete answer.”
↩︎ Turning results into a chunking choice - 3.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-assess-performanceOfficial docs
“Precision is the proportion of retrieved documents that are actually relevant to the user's request.”
↩︎ The retrieval metrics that judge chunking“Recall is the proportion of the ground truth documents that are represented in the retrieved chunks.”
↩︎ The retrieval metrics that judge chunking“Computing recall requires your ground-truth to contain all relevant items.”
↩︎ Exam trap 1“Computing precision does not require knowing all relevant items.”
↩︎ Checkpoint“A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
↩︎ Checkpoint - 4.
“Retrieval quality can be influenced by both the data pipeline (for example, parsing/chunking strategy, metadata strategy, embedding model) and the RAG chain”
↩︎ Is chunking really the cause?“Retrieval quality can be influenced by both the data pipeline (for example, parsing/chunking strategy, metadata strategy, embedding model) and the RAG chain”
↩︎ Exam trap 2 - 5.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-evaluation-monitoring-ragOfficial docs
“it's crucial to evaluate each of the application's components in addition to the application as a whole”
↩︎ Is chunking really the cause?