What you will be able to do
- Name the four broad families of chunking strategy and the LangChain splitter that represents each
- Explain the trade-offs of chunk size and overlap, and how semantic coherence and metadata affect retrieval
- Use the embedding model's token limit and the LLM's context window to bound chunk size
- Explain when parent-child chunking and adding semantic context to chunks are worth trying
Key concept
Chunking as a tunable retrieval knob — No single chunking strategy is correct for every RAG application. Strategy, size, overlap and metadata are settings you pick for your data and then refine by testing, because they decide which text the retriever can return to the LLM.
1.Why chunking is an early quality lever
In a RAG data pipeline, documents are parsed, deduplicated and filtered, then split into chunks. Each chunk is embedded and indexed. The aim is chunks that are small and focused enough to fit in the LLM's context without bringing in distracting text. The retriever can only return chunks, so the way you split documents limits what the LLM ever sees. Databricks calls chunking one of the first layers of optimization in a RAG application.
The cookbook lists five factors to decide when you chunk:
- Strategy: how the text is divided, from sentence, paragraph or token-count splits to document-specific methods. - Size: small chunks pinpoint details but drop the surrounding context. Large chunks keep context but bring in noise. - Overlap: sharing some text between neighbouring chunks keeps continuity, so a fact that falls on a boundary isn't lost. - Semantic coherence: each chunk should hold related information and still make sense on its own. - Metadata: the source document name, section heading or product names help match queries to chunks.
Changing any one of these changes the result set the LLM receives.
Checkpoint 1 of 5· Check yourself
Facts that span two adjacent chunks are often missing from retrieved results. Which chunking factor is designed for this problem?
With overlap, the text near each boundary appears in both neighbouring chunks, so information that falls on a split stays together.
“Overlapping can ensure continuity and context preservation across chunks and improve the retrieval results.”Source: docs.databricks.com
Sources1
2.The four families of chunking strategy
Databricks groups chunking strategies into four families. Moving down the list, each one splits text along more meaningful boundaries and takes more effort to set up. Fixed-size splitting is quick to configure, but it cuts text at arbitrary points, so chunks are often not semantically coherent. Paragraph-based splitting uses the document's own paragraph breaks. Format-specific splitting uses structure such as Markdown or HTML headers. Semantic chunking analyses the content itself, for example with topic modeling, to find where the topic changes.
| Strategy | How boundaries are chosen | Example component | Main caveat |
|---|---|---|---|
| Fixed-size | A set number of characters or tokens | CharacterTextSplitter | Rarely works for production-grade applications |
| Paragraph-based | Natural paragraph boundaries | RecursiveCharacterTextSplitter | Depends on the document having meaningful paragraphs |
| Format-specific | Inherent structure such as Markdown or HTML headers | MarkdownHeaderTextSplitter | Only works for formats that have that structure |
| Semantic | Topic shifts found by analysing the content | SemanticChunker | More involved to set up than basic approaches |
Checkpoint 2 of 5· Match them up
Match each document situation to the chunking strategy family the docs associate with it
Tap a term, then the definition that fits it.
Each family takes its boundaries from a different signal: a count, paragraph breaks, format structure, or changes in topic.
“Formats such as Markdown or HTML have an inherent structure that can define chunk boundaries (for example, markdown headers).”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
A team building a RAG assistant over a corpus of legal contracts is currently using fixed-size, token-count chunking (500 tokens, no overlap). Retrieval evaluation shows that many chunks split individual contract clauses across two consecutive chunks, so the retriever often returns only half of a relevant clause. The contracts have clear, consistent section and clause headers (e.g., "4.2 Termination"). Which change to the chunking strategy is most likely to fix this specific issue?
Correct answer: C — Switch to a document-structure-aware chunking approach that splits on the contract's section and clause headers, so each chunk corresponds to one complete clause.
- A. Shrinking the fixed chunk size to 250 tokens makes it more likely, not less, that a clause is split across a boundary, since the chunker still ignores where clauses actually begin and end.
- B. Retrieving more chunks may happen to pull in the second half of a split clause, but it does not fix the underlying boundary problem and adds noisy, redundant context to every query.
- C. Splitting on the contract's own section and clause headers aligns chunk boundaries with the document's actual structure, so each chunk holds one complete clause instead of a fragment.
- D. Grouping sentences purely by topic similarity ignores the reliable header structure already present in the contracts and can still cut a clause in half if its sentences drift in embedding space.
Sources1
3.Model limits that bound chunk size
Two models put limits on your chunk size.
The embedding model. Every embedding model has a maximum number of input tokens. If a chunk is longer, the model truncates it, and the text past the limit never reaches the vector. For example, bge-large-en-v1.5 accepts at most 512 tokens. A 1,000-token chunk would be embedded from only its first part, even though all of the text is stored.
Not by default. The Databricks GTE example says the model supports up to 8192 tokens, but recommends smaller chunks so the RAG model gets a wider variety of examples. The OpenAI embedding example also uses an 8192-token model and sets max_chunk_tokens = 1024.
The generation LLM. The retriever usually returns several chunks, and all of them have to fit in the LLM's context window together with the prompt. So the right chunk size also depends on how many chunks you plan to retrieve. The embedding model's token limit is a hard maximum. It does not tell you the best chunk size.
Checkpoint 4 of 5· Check yourself
A pipeline embeds 900-token chunks with bge-large-en-v1.5. Retrieval misses facts that appear near the end of long chunks. What is the most likely cause?
bge-large-en-v1.5 has a 512-token limit, and longer chunks are truncated. Text near the end of each chunk is therefore missing from its vector, so queries about it can't match.
“If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”Source: docs.databricks.com
4.Beyond chunk size: semantic context and parent-child chunks
The AI Search retrieval quality guide calls chunk size optimization an active area of research. It recommends that you not spend most of your effort on chunk size. Higher-impact work includes extracting metadata such as entities and topics for filtering, parsing documents well (for example with ai_parse_document), and adding semantic context to chunks. One way to add context is to put the document title, a summary and the section name at the start of each chunk, so the embedding has signal from the whole document.
# Prepend document summary to each chunk
chunk_with_context = f"""
Document: {doc_title}
Summary: {doc_summary}
Section: {section_name}
{chunk_content}
"""The guide lists two advanced strategies that take more effort but can have a bigger impact. Semantic chunking uses embeddings to group sentences by similarity, so related ideas stay together. Parent-child chunking, also called small-to-big retrieval, sidesteps the trade-off between small and large chunks. You search small child chunks, which locate information precisely, and return the larger parent chunk, which gives the LLM context.
Checkpoint 5 of 5· Fill the gap
In this parent-child retrieval sample, which column must be requested so the LLM receives the large context chunk?
# Record child and parent chunks in your source table
for parent_chunk in create_chunks(doc, size=2048): # Large for context
for child_chunk in create_chunks(parent_chunk, size=512): # Small for precision
source_table.append({"text": child_chunk, "parent_text": parent_chunk})
# Search children, return parents
results = index.similarity_search(
query_text="Is attention all you need?",
num_results=10,
columns=["text", " ? "]
)The search matches on the small child text. Requesting the parent_text column returns the 2048-size parent chunk that contains each match.
Source: docs.databricks.comSources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Fixed-size token splitting is an acceptable default for a production RAG application.Why is that wrong?
Fixed-size splitting is quick to set up, but it seldom produces semantically coherent chunks. The cookbook says it rarely works in production.
Covered in The four families of chunking strategy
2.If the embedding model supports 8192 tokens, chunks should be 8192 tokens long.Why is that wrong?
The token limit is a hard maximum, not a target. Databricks still recommends smaller chunks so the LLM receives a wider variety of examples.
Covered in Model limits that bound chunk size
3.Tuning chunk size is the highest-impact way to improve retrieval.Why is that wrong?
The retrieval quality guide recommends spending effort on metadata extraction, high-quality parsing and adding semantic context to chunks instead of over-tuning chunk size.
Covered in Beyond chunk size: semantic context and parent-child chunks
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/quality-data-pipeline-ragOfficial docs
“The choices made on chunking will directly affect the retrieved data the LLM provides”
↩︎ Why chunking is an early quality lever“Smaller chunks may focus on specific details but lose some surrounding contextual information.”
↩︎ Why chunking is an early quality lever“Relevant metadata, such as the source document name, section heading, or product names, can improve retrieval.”
↩︎ Why chunking is an early quality lever“Use the natural paragraph boundaries in the text to define chunks.”
↩︎ The four families of chunking strategy“determine the most appropriate chunk boundaries based on topic shifts”
↩︎ The four families of chunking strategy“For example, bge-large-en-v1.5 has a maximum token limit of 512.”
↩︎ Model limits that bound chunk size“Finding the proper chunking method is both iterative and context-dependent. There is no one-size-fits-all approach.”
↩︎ Key concept“This approach rarely works for production-grade applications.”
↩︎ Exam trap 1“Larger chunks may capture more context but can include irrelevant information or be computationally expensive.”
↩︎ Prediction“Overlapping can ensure continuity and context preservation across chunks and improve the retrieval results.”
↩︎ Checkpoint“Formats such as Markdown or HTML have an inherent structure that can define chunk boundaries (for example, markdown headers).”
↩︎ Checkpoint“If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
↩︎ Checkpoint - 2.
“Context limits: Must fit within LLM context window when retrieving multiple chunks.”
↩︎ Model limits that bound chunk size“Chunk size optimization remains an active area of research.”
↩︎ Beyond chunk size: semantic context and parent-child chunks“Semantic chunking: Group sentences by similarity rather than fixed size.”
↩︎ Beyond chunk size: semantic context and parent-child chunks“Instead of over-optimizing chunk size, focus on:”
↩︎ Exam trap 3 - 3.https://docs.databricks.com/aws/en/ai-search/vector-search-foundation-embedding-model-gte-exampleOfficial docs
“The GTE model supports up to 8192 tokens. However, Databricks recommends that you split the data into smaller context chunks”
↩︎ Model limits that bound chunk size“The GTE model supports up to 8192 tokens. However, Databricks recommends that you split the data into smaller context chunks”
↩︎ Exam trap 2