What you will be able to do
- Place chunking correctly in a RAG data pipeline and name the factors that shape a chunking decision
- Choose between fixed-size, paragraph-based, format-specific and semantic chunking based on how a document is structured
- Size chunks against the embedding model's token limit and the LLM's context window
Key concept
Chunk as a constrained retrieval unit — A chunk is the piece of a document that gets embedded and retrieved. It has to fit within the embedding model's token limit and, along with the other retrieved chunks, within the LLM's context. It should also make sense on its own, so the boundaries need to follow the document's real structure.
1.Where chunking sits, and what you are deciding
A RAG application over unstructured content needs a data pipeline. The pipeline parses raw files, cleans them, cuts them into chunks, embeds each chunk and indexes the vectors. Chunking comes after the text has been extracted and cleaned, and before anything is embedded. Because of where it sits, every later stage works on whatever chunks you produce. The Databricks guidance calls chunking one of the first layers of optimization in a RAG application, because it directly controls what the LLM is given at answer time.
Checkpoint 1 of 6· Put it in order
Put these data pipeline stages in the order the Databricks guidance describes them
- 1.Break the cleaned text into chunks
- 2.Parse the raw data into a more structured format
- 3.Remove duplicate documents
- 4.Filter out unwanted information
Chunking happens only after parsing, deduplication and filtering. Embedding then works on the finished chunks.
“After parsing the raw data into a more structured format, removing duplicates, and filtering out unwanted information, the next step is to break”Source: docs.databricks.com
"Choose a chunk size" is only one part of a chunking decision. The guidance lists five factors. Strategy is the method used to cut the text, for example by sentences, paragraphs, character or token counts, or document-specific rules. Chunk size is a trade-off: smaller chunks focus on specific details but lose surrounding context, and larger chunks keep context but can pull in irrelevant text and cost more to process. Overlap between neighbouring chunks keeps information from being lost at a boundary. Semantic coherence means each chunk should hold related information and still make sense on its own. Metadata, such as the source document name, section heading or product name, gives retrieval queries more to match against.
Sources1
2.Matching the strategy to the document's structure
The Databricks guidance does not name a single best method. It says the right chunk size and method depend on the use case and on the data. It groups the options into four families, and each one ties a kind of document structure to a kind of boundary. Fixed-size splitting ignores structure entirely: it is quick to set up, but it usually does not produce coherent chunks. Paragraph-based splitting uses the paragraph breaks the author already wrote. Format-specific splitting uses explicit markup such as headers. Semantic chunking looks at the content to find where the topic changes, which helps with prose that has little visible structure.
| Strategy | Boundary comes from | Example tool | Trade-off |
|---|---|---|---|
| Fixed-size | A set number of characters or tokens | CharacterTextSplitter | Quick and easy to set up, but chunks are typically not semantically coherent |
| Paragraph-based | Natural paragraph boundaries | RecursiveCharacterTextSplitter | Helps keep related information together |
| Format-specific | Inherent markup such as Markdown headers or HTML sections | MarkdownHeaderTextSplitter, HTML header/section splitters | Only works where the format carries structure |
| Semantic | Topic shifts found by analysing the content | SemanticChunker | More work, but chunks follow natural semantic divisions |
Checkpoint 2 of 6· Match them up
Match each chunking strategy to the LangChain tool the Databricks guidance gives as its example
Tap a term, then the definition that fits it.
Each family pairs with the splitter that finds its kind of boundary: a set count, paragraph breaks, Markdown headers, or topic shifts.
“Tools like LangChain's MarkdownHeaderTextSplitter or HTML header/section-based splitters can be used for this purpose.”Source: docs.databricks.com
Checkpoint 3 of 6· Exam question
A data engineering team ingests thousands of unstructured plain-text support-ticket transcripts and applies a chunking strategy that splits every document into fixed 512-character blocks with no attention to sentence or paragraph boundaries. During evaluation, retrieved chunks frequently cut sentences in half, and the RAG application returns incomplete or confusing context to the LLM. What should the team do to address this?
Correct answer: A — Replace fixed-size chunking with paragraph-based or format-specific chunking that respects natural sentence and section boundaries.
- A. Paragraph-based or format-specific chunking uses the document's natural sentence and section boundaries instead of arbitrary character counts, which produces chunks that stand on their own as coherent, meaningful units. This directly fixes the root cause of chunks cutting sentences in half.
- B. Overlap helps preserve some continuity across chunk boundaries, but it does not stop fixed-size splitting from cutting through sentences at arbitrary character positions in the first place, so the underlying incoherence remains.
- C. Shrinking the fixed chunk size still splits at arbitrary character counts rather than at sentence or paragraph boundaries, so sentences can still be cut, and very small chunks additionally lose surrounding context.
- D. A larger embedding model context window changes how much text can be embedded per call, but it does not change the chunking logic itself, so fixed-size splitting will still cut sentences at the same character boundaries.
Sources1
3.Sizing chunks against the embedding model and the LLM
Two model limits cap chunk size. The first is the embedding model's maximum token limit. The guidance gives bge-large-en-v1.5 as an example, with a limit of 512 tokens. Any chunk longer than the limit is cut off before it is embedded. The second is the LLM's context window. You usually retrieve several chunks for each query, and all of them have to fit into the prompt together. So the limit that matters is chunk size multiplied by the number of chunks you retrieve, not the size of one chunk.
Fitting under the embedding model's limit is necessary but not enough. The AI Search embedding examples point out that the GTE model, and the OpenAI model in the external example, support up to 8192 tokens. Databricks still recommends splitting data into smaller chunks, so that the model generating the answer receives a wider range of examples. The OpenAI example caps chunks at 1024 tokens.
Checkpoint 4 of 6· Fill the gap
This chunking function comes from the external-embedding-model example. The model supports 8192 tokens, but the example follows Databricks' advice to use smaller chunks. Which value fills the blank?
max_chunk_tokens = ?
encoding = tiktoken.get_encoding("cl100k_base")
def chunk_text(text):
# Encode and then decode within the UDF
tokens = encoding.encode(text)
chunks = []
while tokens:
chunk_tokens = tokens[:max_chunk_tokens]
chunk_text = encoding.decode(chunk_tokens)
chunks.append(chunk_text)
tokens = tokens[max_chunk_tokens:]
return chunksThe example sets 1024 tokens, well under the model's 8192-token limit. This splits by token count only, so it is the fixed-size strategy.
Source: docs.databricks.comFor choosing a size within those limits, the AI Search retrieval quality guide suggests three starting points to test. It also says chunk size optimization is still an active area of research, so treat these as experiments rather than answers.
# Common configurations to test
small_chunks = 256 # Better for precise fact retrieval
medium_chunks = 512 # Balanced approach
large_chunks = 1024 # More context per chunkCheckpoint 5 of 6· Check yourself
A RAG chain retrieves ten chunks for each question. You are thinking of moving from 512-token to 1024-token chunks. Which constraint does the guidance say you must check?
Every retrieved chunk goes into the prompt, so the total size of all of them has to fit the LLM's context window.
“Context limits: Must fit within LLM context window when retrieving multiple chunks.”Source: docs.databricks.com
Checkpoint 6 of 6· Exam question
An engineer builds a RAG pipeline over long onboarding guides. During testing, they notice that a key instruction spanning the end of one chunk and the beginning of the next is never retrieved in full: the retriever returns one chunk or the other, but neither chunk alone contains the complete instruction. Which adjustment to the chunking configuration most directly addresses this failure?
Correct answer: A — Configure adjacent chunks to overlap by a set number of tokens so boundary content appears in more than one chunk.
- A. Adding overlap between adjacent chunks means content near a boundary is duplicated into both chunks, so an instruction that straddles the boundary will appear in full inside at least one chunk and can be retrieved intact.
- B. Returning more candidate chunks increases the chance both neighboring chunks are retrieved together, but neither chunk individually contains the complete instruction, so the response context is still fragmented across two partial chunks.
- C. A domain-tuned embedding model can improve semantic matching for retrieval, but it does not change where chunk boundaries fall, so the instruction would still be split across two chunks.
- D. Physical storage order in the vector index does not affect which chunks are returned for a similarity search, so this does not help the retriever surface the complete instruction.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Fixed-size chunking by character or token count is a sound default for a production RAG system.Why is that wrong?
Fixed-size splitting is quick to set up, but it typically does not produce semantically coherent chunks. The guidance says it rarely works in production. Use the document's paragraphs, markup or topics to set boundaries instead.
Covered in Matching the strategy to the document's structure
2.If the embedding model accepts 8192 tokens, chunks should be close to 8192 tokens.Why is that wrong?
The limit is a ceiling, not a target. Databricks recommends smaller chunks so the reasoning model receives a wider range of examples.
Covered in Sizing chunks against the embedding model and the LLM
3.A chunk longer than the embedding model's limit causes an error, so you will notice it.Why is that wrong?
Oversized chunks are truncated without any error, so the text past the limit never reaches the embedding.
Covered in Sizing chunks against the embedding model and the LLM
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/quality-data-pipeline-ragOfficial docs
“Smaller chunks may focus on specific details but lose some surrounding contextual information.”
↩︎ Where chunking sits, and what you are deciding“aim to create semantically coherent chunks that contain related information but can stand independently as a meaningful unit of text”
↩︎ Where chunking sits, and what you are deciding“Finding the proper chunking method is both iterative and context-dependent. There is no one-size-fits-all approach.”
↩︎ Matching the strategy to the document's structure“Use the natural paragraph boundaries in the text to define chunks.”
↩︎ Matching the strategy to the document's structure“For example, bge-large-en-v1.5 has a maximum token limit of 512.”
↩︎ Sizing chunks against the embedding model and the LLM“Segmenting large documents into smaller, semantically concentrated chunks ensures that retrieved data fits in the LLM's context”
↩︎ Key concept“This approach rarely works for production-grade applications.”
↩︎ Exam trap 1“If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
↩︎ Exam trap 3“After parsing the raw data into a more structured format, removing duplicates, and filtering out unwanted information, the next step is to break”
↩︎ Checkpoint“Formats such as Markdown or HTML have an inherent structure that can define chunk boundaries”
↩︎ Prediction“Tools like LangChain's MarkdownHeaderTextSplitter or HTML header/section-based splitters can be used for this purpose.”
↩︎ Checkpoint“If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
↩︎ Prediction - 2.
“Chunk size optimization remains an active area of research.”
↩︎ Sizing chunks against the embedding model and the LLM“Context limits: Must fit within LLM context window when retrieving multiple chunks.”
↩︎ Checkpoint
Also cited
- https://docs.databricks.com/aws/en/ai-search/vector-search-foundation-embedding-model-gte-exampleOfficial docs
“The GTE model supports up to 8192 tokens. However, Databricks recommends that you split the data into smaller context chunks”
↩︎ Exam trap 2