What you will be able to do
- Place chunking correctly in a RAG data pipeline and name the five factors that shape a chunking design
- Choose between fixed-size, paragraph-based, format-specific and semantic chunking for a given document structure
- Size chunks against embedding-model token limits and the LLM context window, and know which optimizations outrank chunk-size tuning
Key concept
Semantically coherent chunk — A chunk is the unit your retriever returns, so it should hold related information and still make sense on its own. Every chunking strategy and size choice is a way of getting closer to that while staying inside model limits.
1.Where chunking sits and what it decides
A RAG data pipeline turns raw documents into something a retriever can query. The Databricks unstructured-pipeline guide breaks it into these stages: choosing the corpus, preprocessing (parsing, enrichment, metadata extraction, deduplication, filtering), chunking, embedding, and indexing. Chunking runs once the text is clean and before anything is embedded. This is where you decide the unit of text that retrieval will hand to the LLM.
Chunks have two jobs. They must be small enough that several of them fit in the LLM's context together, and focused enough that a retrieved chunk isn't padded with off-topic text. The guide puts it as chunking "ensures that retrieved data fits in the LLM's context while minimizing the inclusion of distracting or irrelevant information." Everything downstream sees only what the chunks contain, so the guide calls chunking "one of the first layers of optimization in an RAG application."
Checkpoint 1 of 5· Put it in order
Put these data-pipeline stages in the order the Databricks guide describes
- 1.Parse the raw documents into a more structured format
- 2.Break the cleaned documents into chunks
- 3.Convert each chunk into a vector with an embedding model
- 4.Remove duplicates and filter out unwanted documents
Chunking comes after parsing, deduplication and filtering, and embedding comes after chunking. Chunk text that hasn't been cleaned would carry the noise forward into every vector.
“removing duplicates, and filtering out unwanted information, the next step is to break it down into smaller, manageable units called chunks”Source: docs.databricks.com
The guide names five factors to weigh when you design chunking:
- Chunking strategy: how the text is divided (by sentence, paragraph, character or token count, or document-specific structure). - Chunk size: small chunks focus on specific details but lose surrounding context. Large chunks keep context but bring in irrelevant material and cost more to compute. - Overlap: repeating some text between adjacent chunks so that information sitting on a boundary isn't lost. - Semantic coherence: chunks that can stand on their own, built by following paragraphs, sections or topic boundaries. - Metadata: source document name, section heading or product names attached to each chunk to help match queries.
Chunking isn't a self-contained problem. The quality overview points out that retrieval quality depends on choices in both the pipeline and the chain, and the pipeline choices it lists include the chunking strategy.
2.Four strategies, matched to document structure
The guide says outright that "There is no one-size-fits-all approach." The right method depends on your use case and on how your documents are structured. It groups the options into four families, from cheapest to most involved.
Fixed-size chunking splits text every N characters or tokens. It's fast to set up, but the cuts land wherever the count runs out, often mid-sentence or mid-table, so the chunks are rarely coherent. The guide's verdict: "This approach rarely works for production-grade applications."
Paragraph-based chunking follows the paragraph breaks the author already wrote. Paragraphs usually hold related information, so this preserves coherence at little cost.
Format-specific chunking uses the structure built into formats like Markdown or HTML. Headers or sections become the chunk boundaries, so a chunk lines up with a section a human would recognise.
Semantic chunking looks at the content to find where the topic changes, for example with topic modeling or by comparing sentence embeddings. The AI Search retrieval quality guide describes it as grouping sentences by similarity instead of by fixed size, which keeps related ideas together. It takes more effort, but the boundaries match the text's natural divisions.
| Strategy | How boundaries are chosen | Example splitter | Trade-off |
|---|---|---|---|
| Fixed-size | A predetermined number of characters or tokens | LangChain CharacterTextSplitter | Quick and easy, but typically not semantically coherent |
| Paragraph-based | Natural paragraph boundaries in the text | LangChain RecursiveCharacterTextSplitter | Helps preserve coherence, since paragraphs often hold related information |
| Format-specific | Inherent structure such as Markdown headers or HTML sections | LangChain MarkdownHeaderTextSplitter | Only available where the format carries structure |
| Semantic | Topic shifts in the content | LangChain SemanticChunker | More involved, but aligned with natural semantic divisions |
Checkpoint 2 of 5· Match them up
Match each chunking strategy to how it sets chunk boundaries
Tap a term, then the definition that fits it.
Each strategy gets its boundaries from a different place: an arbitrary count, the author's paragraphs, the format's markup, or an analysis of the content itself.
“Use the natural paragraph boundaries in the text to define chunks.”Source: docs.databricks.com
Whichever strategy you pick, overlap is a separate setting. Repeating a little text between neighbouring chunks protects facts that would otherwise be cut in two. According to the guide, overlap can preserve continuity and context and improve retrieval results.
Checkpoint 3 of 5· Exam question
A team is building a RAG chatbot over internal engineering documentation written in Markdown, where every page uses consistent H1/H2/H3 headers to organize topics such as installation, configuration, and troubleshooting. Which chunking strategy is best suited to this document set to preserve topical boundaries?
Correct answer: C — Split each markdown file at its header boundaries and keep the header hierarchy attached to each chunk.
- A. Fixed-size chunking ignores the document's structure entirely, so a 500-character boundary can fall in the middle of a configuration step or split a header from the content it introduces, which loses the topical grouping the headers already provide.
- B. Semantic chunking via topic modeling is useful when a document lacks structural markers, but deliberately ignoring existing, reliable headers discards a cheaper and more precise signal for where topics change, making this unnecessarily complex for this document set.
- C. Document-structure-aware chunking that splits at header boundaries and keeps the header hierarchy attached directly matches the way this documentation is already organized, producing chunks that stay topically coherent and carry useful context like section titles.
- D. Splitting only on paragraph breaks with a plain-text length target treats the markdown the same as unstructured prose, throwing away the header hierarchy that would otherwise cleanly separate installation, configuration, and troubleshooting content.
3.Sizing chunks against embedding and LLM limits
Two limits cap chunk size. The first is the embedding model's maximum token count. The guide's example is bge-large-en-v1.5, which has a limit of 512 tokens, and chunks longer than that get truncated. The second is the LLM context window. The retrieval quality guide notes that retrieved chunks must fit inside it, and you usually retrieve several at a time, so the space is shared.
A large embedding limit doesn't mean your chunks should be that large. The AI Search example notebooks use a model that accepts 8,192 tokens, yet they still advise splitting into smaller chunks so the reasoning model gets a wider variety of examples. The OpenAI-embedding example splits on a fixed token count of 1,024:
def chunk_text(text):
# Encode and then decode within the UDF
tokens = encoding.encode(text)
chunks = []
while tokens:
chunk_tokens = tokens[:max_chunk_tokens]
chunk_text = encoding.decode(chunk_tokens)
chunks.append(chunk_text)
tokens = tokens[max_chunk_tokens:]
return chunksFor experiments, the retrieval quality guide suggests testing 256 tokens for precise fact retrieval, 512 as a middle ground, and 1,024 for more context per chunk. Smaller chunks pinpoint information better but can lose context. Larger chunks keep context but make the relevant passage harder to locate.
The guide also warns against spending too long on size. It calls chunk size an active research area, then says: "Instead of over-optimizing chunk size, focus on:" extracting metadata (entities, topics, categories) for filtering, parsing documents cleanly, and adding semantic metadata such as summaries and section headers to chunks.
Checkpoint 4 of 5· Check yourself
A team has spent two sprints trying 256, 512 and 1,024-token chunks, and the gains are small. What does the AI Search retrieval quality guide say to focus on instead?
The guide lists metadata extraction, high-quality parsing and semantic metadata as more impactful than further chunk-size tuning.
“Information extraction for metadata: Extract entities, topics, and categories to enable precise filtering.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A support knowledge base consists of long compiled FAQ articles where a single page can cover several unrelated subtopics with no headers or paragraph markers to indicate where one topic ends and another begins. Which chunking approach directly addresses this document characteristic?
Correct answer: C — Embed consecutive sentences and start a new chunk where similarity between neighbors drops sharply.
- A. A fixed token count applied from the start of the document has no awareness of where subtopics change, so it will frequently cut a chunk in the middle of one subtopic or merge the end of one subtopic with the start of another.
- B. Adding overlap helps preserve continuity across an already-chosen boundary, but the underlying split points are still arbitrary character counts, so this does not solve the problem of subtopics shifting mid-article without any structural cue.
- C. Semantic chunking based on embedding similarity between adjacent sentences is designed for exactly this situation: it finds natural topic shifts even when there are no headers or paragraph markers to rely on, producing chunks aligned with actual subtopic boundaries.
- D. These FAQ articles have no headers or paragraph markers, so relying on HTML heading tags to define split points will not work if the underlying markup does not consistently tag each subtopic with its own heading.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Fixed-size character or token splitting is a sound production default because it's simple and predictable.Why is that wrong?
Arbitrary character or token counts usually produce chunks that aren't semantically coherent. The guide says this approach rarely works in production; prefer paragraph, format-specific or semantic boundaries.
2.Bigger chunks are always safer because they keep more context.Why is that wrong?
Chunks over the embedding model's token limit are truncated, and retrieved chunks must also fit in the LLM context window together. Bigger chunks keep context but make relevant information harder to pinpoint.
3.Tuning chunk size is the main lever for improving retrieval quality.Why is that wrong?
The retrieval quality guide calls chunk size an active research area and points to metadata extraction, high-quality parsing and semantic metadata as more impactful.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/quality-data-pipeline-ragOfficial docs
“ensures that retrieved data fits in the LLM's context while minimizing the inclusion of distracting or irrelevant information”
↩︎ Where chunking sits and what it decides“making it one of the first layers of optimization in an RAG application”
↩︎ Where chunking sits and what it decides“There is no one-size-fits-all approach.”
↩︎ Four strategies, matched to document structure“Formats such as Markdown or HTML have an inherent structure that can define chunk boundaries”
↩︎ Four strategies, matched to document structure“Overlapping can ensure continuity and context preservation across chunks and improve the retrieval results.”
↩︎ Four strategies, matched to document structure“For example, bge-large-en-v1.5 has a maximum token limit of 512.”
↩︎ Sizing chunks against embedding and LLM limits“aim to create semantically coherent chunks that contain related information but can stand independently as a meaningful unit of text”
↩︎ Key concept“This approach rarely works for production-grade applications.”
↩︎ Exam trap 1“removing duplicates, and filtering out unwanted information, the next step is to break it down into smaller, manageable units called chunks”
↩︎ Checkpoint“Use the natural paragraph boundaries in the text to define chunks.”
↩︎ Checkpoint“If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
↩︎ Prediction - 2.
“Retrieval quality can be influenced by both the data pipeline (for example, parsing/chunking strategy, metadata strategy, embedding model)”
↩︎ Where chunking sits and what it decides - 3.
“Use embeddings to find natural semantic boundaries.”
↩︎ Four strategies, matched to document structure“Smaller chunks: Better localization of specific information, but may lose context.”
↩︎ Sizing chunks against embedding and LLM limits“Context limits: Must fit within LLM context window when retrieving multiple chunks.”
↩︎ Exam trap 2“Instead of over-optimizing chunk size, focus on:”
↩︎ Exam trap 3“Information extraction for metadata: Extract entities, topics, and categories to enable precise filtering.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/ai-search/vector-search-external-embedding-model-exampleOfficial docs
“split the data into smaller context chunks so that you can feed a wider variety of examples into the reasoning model”
↩︎ Sizing chunks against embedding and LLM limits