What you will be able to do
- Explain why Databricks advises against over-optimizing chunk size, and what to improve instead
- Design parent-child (small-to-big) chunking that searches small chunks and returns larger ones
- Use ai_parse_document structure and ai_prep_search to build chunks that carry document context
1.Beyond chunk size: semantic boundaries
The AI Search retrieval quality guide pushes back on tuning chunk size alone. It notes that recent research shows embeddings can fail to capture basic information in long contexts, which makes size a subtle choice rather than a dial to max out. It then names areas that usually matter more: extracting entities, topics and categories as metadata for filtering; high-quality parsing, for which it recommends ai_parse_document; and adding semantic metadata such as document summaries and section headers to chunks. Poor parsing, such as missing tables or broken formatting, directly lowers retrieval quality, and no chunk size fixes that.
Its first advanced technique is semantic chunking. Instead of cutting at a fixed size, it groups sentences by similarity and uses embeddings to find natural semantic boundaries. The result keeps related ideas together and preserves context better. It needs more work than a splitter with a size setting, and the guide puts it among techniques that can have a bigger impact.
Checkpoint 1 of 5· Check yourself
How does semantic chunking, as the AI Search retrieval quality guide describes it, decide where one chunk ends?
Semantic chunking sets boundaries from the content's meaning, measured with embeddings, rather than from a count or from markup.
“Semantic chunking: Group sentences by similarity rather than fixed size.”Source: docs.databricks.com
Sources1
2.Parent-child chunking: search small, return big
The size trade-off from basic chunking still applies: small chunks pinpoint facts, and large chunks keep context. Parent-child chunking, also called small-to-big retrieval, avoids having to choose. Each document is cut into large parent chunks, and each parent is cut again into small child chunks. Both are stored on the same row of the source table. Similarity search runs against the precise child text, and the query also returns the parent text, so the LLM gets the surrounding context.
# Record child and parent chunks in your source table
for parent_chunk in create_chunks(doc, size=2048): # Large for context
for child_chunk in create_chunks(parent_chunk, size=512): # Small for precision
source_table.append({"text": child_chunk, "parent_text": parent_chunk})
# Search children, return parents
results = index.similarity_search(
query_text="Is attention all you need?",
num_results=10,
columns=["text", "parent_text"]
)Embed and search the child text column. The 512-token children give the embedding a narrow, precise focus. The 2048-token parent is only returned as context, so it never has to be precise enough to match a query.
Checkpoint 2 of 5· Check yourself
A team needs precise matching on specific clauses in long contracts, but the LLM also needs the surrounding section to answer correctly. Which design fits the guidance?
Small-to-big retrieval matches on small chunks for precision and passes the larger parent to the LLM for context, which gives you both.
“# Search children, return parents”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
A team is preparing a corpus of Markdown product documentation, where each file uses `#`, `##`, and `###` headers to organize content into sections and subsections. They want chunks that align with the document's logical structure rather than arbitrary character counts, so that each chunk corresponds to a coherent section. Which chunking approach best fits this requirement?
Correct answer: A — A format-specific approach that splits at Markdown header boundaries, keeping each section intact via a header-aware splitter.
- A. Format-specific chunking uses the document's inherent structure, such as Markdown header levels, to define chunk boundaries, so each chunk maps directly to a section or subsection instead of an arbitrary span of text.
- B. Splitting into equal character blocks ignores the Markdown headers entirely, so a section could be split mid-way while an unrelated section boundary falls inside a chunk, defeating the goal of section-aligned chunks.
- C. Topic modeling is useful when a document lacks explicit structural markers, but here the Markdown headers already provide reliable section boundaries, so relying on inferred topic shifts is unnecessary and less precise than using the existing structure.
- D. Sampling a subset of sentences discards content rather than defining chunk boundaries, so it would not produce complete, retrievable chunks that represent each section.
Sources1
3.Structure-aware chunks that carry their own context
A chunk taken out of a long report can lose information that only appeared elsewhere, such as which company, which section or which year it belongs to. Structure-aware chunking needs that structure to be detected first. ai_parse_document breaks a document into typed elements, including section_header, page_header, page_footer and page_number. Those types let you split at section boundaries and handle repeated page furniture separately from body text. The retrieval quality guide's data-cleaning advice says to remove boilerplate such as headers, footers and page numbers, preserve document structure, and keep semantic boundaries intact when chunking.
The guide's simplest form of enrichment adds document-level context to the start of each chunk before embedding. This gives the embedding model extra semantic signal and helps with queries about document-level ideas.
# Prepend document summary to each chunk
chunk_with_context = f"""
Document: {doc_title}
Summary: {doc_summary}
Section: {section_name}
{chunk_content}
"""Databricks also has a built-in function for this, ai_prep_search (Beta, Databricks Runtime 18.2 or above). It takes the structured output of ai_parse_document, or plain text or markdown. It splits the content into semantic chunks and enriches each chunk with document-level context. Each output chunk has two text fields: chunk_to_retrieve, the raw chunk text, and chunk_to_embed, an enriched string with the document title, page header and footer, section header, caption, footnote and page number. It can also include LLM-discovered or schema-defined document fields, a one-sentence summary of the document and, for tables, a summary plus related questions. When the flattened output is written to a Delta table for an AI Search index, chunk_to_embed is the embedding column and chunk_id is the primary key.
Checkpoint 4 of 5· Fill the gap
This example produces search-ready chunks from raw files in a Unity Catalog volume. Which function fills the blank?
WITH parsed_documents AS (
SELECT ai_parse_document(content) AS parsed
FROM READ_FILES('/Volumes/mydata/documents/', format => 'binaryFile')
)
SELECT ? (parsed) AS result
FROM parsed_documents;ai_parse_document handles the parsing in the CTE. ai_prep_search then turns the parsed VARIANT into enriched semantic chunks.
Source: docs.databricks.comCheckpoint 5 of 5· Check yourself
In ai_prep_search output, which field contains the raw chunk text with no added document context?
chunk_to_embed is the context-enriched string used for embedding. Its Content part is the same as chunk_to_retrieve, which holds the raw text.
“Content: the raw chunk text. Same value as the chunk_to_retrieve field for the chunk.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Finding the best chunk size is the most effective way to improve RAG retrieval.Why is that wrong?
The guide warns against over-optimizing chunk size. It points instead to metadata extraction, high-quality parsing, and adding summaries and section headers to chunks.
Covered in Beyond chunk size: semantic boundaries
2.ai_prep_search's chunk_to_embed field is just the raw chunk text, so it does not matter which field you index.Why is that wrong?
chunk_to_embed adds document-level context to the raw text. That is why it is the embedding column, with chunk_id as the primary key.
Covered in Structure-aware chunks that carry their own context
Practise it for real
Turn raw documents in a Unity Catalog volume into one row per context-enriched chunk, ready to use as an AI Search source table.
1.In a notebook or the SQL editor on Databricks Runtime 18.2 or above, run ai_parse_document over READ_FILES('/Volumes/mydata/documents/', format => 'binaryFile') in a CTE named parsed_documents, replacing the path with your own volume.
Why: ai_prep_search works best on structured parser output, which marks section headers, page headers and tables.
You should see: One row per file, with a VARIANT column called parsed.
2.Wrap the parsed column in ai_prep_search with map('schema', '["company", "document_type", "fiscal_year"]') and alias the result as result.
Why: The schema option sets which document-level fields are extracted and attached to every chunk.
You should see: One row per document, with result:document.contents holding an array of chunks.
3.Use LATERAL variant_explode(prepped_documents.result:document.contents) AS chunk to flatten the array, and select chunk.value:chunk_id::STRING and chunk.value:chunk_to_embed::STRING along with the metadata keys.
Why: A vector index source table needs one row per chunk, with a primary key and a text column to embed.
You should see: One row per chunk. chunk_to_embed starts with 'Document Title:' followed by the section header and page-number lines.
4.Compare chunk_to_embed with chunk_to_retrieve for a chunk that contains a table.
Why: This shows the enrichment that only appears for tables.
You should see: chunk_to_embed includes a 'Table summary:' line and a 'Related questions:' block that chunk_to_retrieve does not have.
Stuck? Get a nudge
If the call fails on serverless compute, check that the environment version is 3 or above, since this function needs the VARIANT type.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Poor parsing (missing tables, broken formatting) directly impacts retrieval quality.”
↩︎ Beyond chunk size: semantic boundaries“Semantic metadata: Add document summaries and section headers to chunks.”
↩︎ Beyond chunk size: semantic boundaries“Parent-child chunking (small-to-big retrieval):”
↩︎ Parent-child chunking: search small, return big“Remove boilerplate (headers, footers, page numbers).”
↩︎ Structure-aware chunks that carry their own context“Provides additional semantic signal for embedding models.”
↩︎ Structure-aware chunks that carry their own context“Instead of over-optimizing chunk size, focus on:”
↩︎ Exam trap 1“Semantic chunking: Group sentences by similarity rather than fixed size.”
↩︎ Checkpoint“# Search children, return parents”
↩︎ Checkpoint - 2.
“section_header: A heading or subheading that denotes the start of a section.”
↩︎ Structure-aware chunks that carry their own context - 3.
“the function splits content into semantic chunks, enriches each chunk with document-level context”
↩︎ Structure-aware chunks that carry their own context“using chunk_to_embed as the embedding column and chunk_id as the primary key”
↩︎ Exam trap 2“Content: the raw chunk text. Same value as the chunk_to_retrieve field for the chunk.”
↩︎ Checkpoint