What you will be able to do
- Clean source text so that chunk boundaries follow content rather than page furniture
- Design parent-child (small-to-big) chunking that searches precise chunks and returns wider context
- Add semantic context to chunks inline or as metadata columns, and know what each option requires downstream
- Read the chunk_to_embed and chunk_to_retrieve output of ai_prep_search and know which column feeds the index
1.Clean the text so chunks carry content, not boilerplate
A chunking strategy can only be as good as the text it receives. Page headers, footers and page numbers repeat on every page of a PDF. If they reach the chunker, they end up inside chunks, where they add noise to the embedding and can break a sentence across a boundary. The unstructured-pipeline guide tells you to "Preprocess the extracted text to remove irrelevant or noisy information, such as headers, footers, or special characters."
The AI Search retrieval quality guide pairs cleaning with chunking in three rules: remove boilerplate (headers, footers, page numbers), preserve document structure (headings, lists, tables), and maintain semantic boundaries when chunking. The second rule matters for the advanced designs on this page. Structure-aware chunking and section-header metadata only work if the headings survive cleaning.
Parsing comes before all of this. For PDFs and complex documents, Databricks recommends ai_parse_document, because poor parsing (missing tables, broken formatting) directly hurts retrieval quality.
Checkpoint 1 of 6· Check yourself
Which preprocessing step goes AGAINST the retrieval quality guide's cleaning advice?
Boilerplate should go, but document structure should be kept. Headings, lists and tables are what structure-aware chunking and section metadata depend on.
“Preserve document structure (headings, lists, tables).”Source: docs.databricks.com
2.Parent-child chunking: search small, return big
The previous page described a size trade-off: small chunks locate facts precisely, while large chunks keep context. Parent-child chunking uses both instead of compromising between them. You cut each document into large parent chunks, then cut each parent into small child chunks. Each child row stores its own text plus its parent's text. The index embeds and searches the children, and the query asks for the parent column back, so the LLM gets the surrounding context the small match lacked.
The retrieval quality guide lists this among the advanced approaches that take more effort but can have a bigger impact, alongside semantic chunking:
# Record child and parent chunks in your source table
for parent_chunk in create_chunks(doc, size=2048): # Large for context
for child_chunk in create_chunks(parent_chunk, size=512): # Small for precision
source_table.append({"text": child_chunk, "parent_text": parent_chunk})
# Search children, return parents
results = index.similarity_search(
query_text="Is attention all you need?",
num_results=10,
columns=["text", "parent_text"]
)Parent text is what gets returned to the LLM, and with num_results=10 up to ten parents may arrive together. They all have to fit in the context window with the prompt. The embedding limit applies to the 512-token children, and the context-window limit applies to the parents.
Checkpoint 2 of 6· Check yourself
In the parent-child code sample, what happens at query time?
The query matches the small children, and the parent_text column requested in the results supplies the larger context.
“# Search children, return parents”Source: docs.databricks.com
Sources2
3.Giving chunks document-level context
Once a chunk is cut out of its document, it often loses the context that made it findable. A paragraph about "adjustment procedures" doesn't say which manual or which system it belongs to. The retrieval quality guide recommends adding semantic context because it gives embedding models an extra semantic signal, gives rerankers more to score on, and helps with queries that refer to document-level concepts. It offers two options.
Option 1: put the context inside the chunk. Prepend the document title, summary and section name to the chunk text before embedding:
# Prepend document summary to each chunk
chunk_with_context = f"""
Document: {doc_title}
Summary: {doc_summary}
Section: {section_name}
{chunk_content}
"""Option 2: store the context in separate metadata columns. The chunk text stays as it is, and the summary, section and keywords sit beside it:
# Store semantic metadata for reranker to use
metadata = {
"doc_summary": "Technical manual for brake system maintenance",
"section": "Emergency brake adjustment procedures",
"keywords": ["brake", "safety", "adjustment"]
}The two options differ in a way that matters. Option 1 changes the embedded text, so the context affects the vector directly. Option 2 doesn't touch the vector, and the guide warns that it "requires downstream processing to leverage the metadata." For semantic metadata, that means reranking with the columns_to_rerank parameter. For keyword-only metadata, it means hybrid (full-text) search against those fields. The unstructured-pipeline guide makes the general point too: store metadata alongside chunks or their embeddings, because it narrows down what gets retrieved.
Checkpoint 3 of 6· Check yourself
A team stores doc_summary and section as separate metadata columns next to each chunk, with no other changes. Retrieval doesn't improve. What is missing?
Separate columns don't change the vectors. They only help if a reranker or keyword search is configured to use them.
“For semantic metadata: Use reranking with columns_to_rerank parameter to consider these columns.”Source: docs.databricks.com
Checkpoint 4 of 6· Exam question
During evaluation, a RAG engineer notices that answers referencing information near a chunk boundary are frequently incomplete, because the fact needed to answer the question is split across two adjacent chunks. Which change to the chunking configuration best addresses this specific problem?
Correct answer: B — Configure the chunker to carry a window of trailing sentences into the start of the next chunk.
- A. Retrieving more chunks increases the chance that a second chunk with the missing fact is also returned, but it does not guarantee it, and it adds noise and cost to the prompt without directly fixing the boundary-splitting issue.
- B. Carrying a window of trailing sentences from one chunk into the next, known as chunk overlap, ensures continuity across the boundary so a fact split between two chunks is likely to appear intact in at least one of them.
- C. Switching to a fixed-character splitter changes where boundaries fall but does not add continuity across them, so facts can still be split between chunks just as before, only at different points in the text.
- D. Shrinking the chunk size increases the number of boundaries in the document, which makes it more likely, not less, that a given fact will end up split across two separate chunks.
4.ai_prep_search: context-enriched chunks in SQL
Databricks provides a SQL function that applies Option 1 for you. ai_prep_search takes the output of ai_parse_document and returns chunks. Each chunk comes in two versions: chunk_to_retrieve, the raw chunk text, and chunk_to_embed, the same text wrapped in labelled document context. The documentation states that the Content field of chunk_to_embed is "the raw chunk text. Same value as the chunk_to_retrieve field for the chunk." The template looks like this:
Document Title: {doc_title}
Page Header: {page_header}
Page Footer: {page_footer}
Section Header: {section_header}
Caption: {caption}
Footnote: {footnote}
Page Number: {page_number}
{additional_document_fields}
{document_context_sentence}
Table summary: {table_summary}
Content:
{chunk_to_retrieve}
Related questions:
{qa_text}Three details are worth knowing. The fixed fields (title, header, footer, section header and so on) always appear with their label, left empty when there's no value. For chunks that contain a table, the function adds an LLM-written table summary and related natural-language questions, which are meant to improve recall for table content. Without the schema option, an LLM decides which document fields to extract for each document. Passing schema makes every chunk carry the same keys. The flattened output becomes a Delta table that can feed an AI Search index:
| Column | Type | Role |
|---|---|---|
| chunk_id | STRING | Primary key of the AI Search index |
| chunk_to_embed | STRING | Embedding column: the chunk wrapped in labelled document context |
| chunk_to_retrieve | STRING | The raw chunk text |
| metadata | VARIANT | Extracted document fields, such as the keys passed in the schema option |
| source_uri | STRING | Path of the source file read from the volume |
Checkpoint 5 of 6· Fill the gap
Which function turns the parsed documents into chunks in this pipeline?
prepped_documents AS (
SELECT
path,
? (parsed) AS result
FROM parsed_documents
)
SELECT
chunk.value:chunk_id::STRING AS chunk_id,
chunk.value:chunk_position::INT AS chunk_position,
chunk.value:chunk_to_retrieve::STRING AS chunk_to_retrieve,
chunk.value:chunk_to_embed::STRING AS chunk_to_embed,
chunk.value:metadata AS metadata,
prepped_documents.path AS source_uri
FROM
prepped_documents,
LATERAL variant_explode(prepped_documents.result:document.contents) AS chunk;ai_parse_document runs earlier, on the raw binary content. ai_prep_search takes its parsed output and produces the chunks, which variant_explode then flattens into one row per chunk.
Source: docs.databricks.comCheckpoint 6 of 6· Check yourself
You're creating an AI Search index from the flattened ai_prep_search table. Which column should be the embedding column?
The documentation names chunk_to_embed as the embedding column and chunk_id as the primary key. chunk_to_retrieve holds the raw text without the added context.
“using chunk_to_embed as the embedding column and chunk_id as the primary key”Source: docs.databricks.com
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.In parent-child retrieval, the large parent chunks are embedded and searched because they carry more context.Why is that wrong?
The small children are what get embedded and searched, for precision. The parent text is stored next to each child and returned as context.
2.Storing a document summary in a separate metadata column improves vector retrieval on its own.Why is that wrong?
Separate metadata columns don't change the embedded text. They only help when a reranker (columns_to_rerank) or hybrid keyword search is set up to use them.
Covered in Giving chunks document-level context
Practise it for real
Turn documents in a Unity Catalog volume into context-enriched chunk rows with ai_parse_document and ai_prep_search, then compare what gets embedded with what gets retrieved.
1.Run a query that reads your volume with READ_FILES('/Volumes/mydata/documents/', format => 'binaryFile') and selects path plus ai_parse_document(content) AS parsed, wrapped in a parsed_documents CTE.
Why: ai_prep_search expects parsed documents, and clean parsing comes before any chunking.
You should see: One row per file, with a parsed result for each document.
2.Add a prepped_documents CTE that selects path and ai_prep_search(parsed) AS result from parsed_documents.
Why: This step does the chunking and builds the context-enriched chunk_to_embed text.
You should see: One result per document, holding a document.contents array of chunks.
3.Flatten with LATERAL variant_explode(prepped_documents.result:document.contents) AS chunk, and project chunk_id, chunk_position, chunk_to_retrieve, chunk_to_embed, metadata and path AS source_uri.
Why: An AI Search source table needs one row per chunk, with a primary key and an embedding column.
You should see: Rows with the columns chunk_id STRING, chunk_position INT, chunk_to_retrieve STRING, chunk_to_embed STRING, metadata VARIANT, source_uri STRING.
4.Compare chunk_to_embed and chunk_to_retrieve for a few rows, including a chunk that contains a table.
Why: Seeing the labelled header fields, context sentence and table aids shows exactly what extra signal the embedding gets.
You should see: chunk_to_embed starts with labelled fields such as Document Title and Section Header, with the raw chunk_to_retrieve text under Content. Table chunks also carry a table summary and related questions.
Stuck? Get a nudge
If the metadata fields vary from document to document, rerun ai_prep_search with map('schema', '["company", "document_type", "fiscal_year"]') so that every chunk carries the same keys.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/quality-data-pipeline-ragOfficial docs
“Preprocess the extracted text to remove irrelevant or noisy information, such as headers, footers, or special characters.”
↩︎ Clean the text so chunks carry content, not boilerplate“Storing metadata alongside chunked documents or their corresponding embeddings is essential for optimal performance.”
↩︎ Giving chunks document-level context - 2.
“Remove boilerplate (headers, footers, page numbers).”
↩︎ Clean the text so chunks carry content, not boilerplate“Poor parsing (missing tables, broken formatting) directly impacts retrieval quality.”
↩︎ Clean the text so chunks carry content, not boilerplate“Parent-child chunking (small-to-big retrieval):”
↩︎ Parent-child chunking: search small, return big“These techniques require more effort but can have a bigger impact:”
↩︎ Parent-child chunking: search small, return big“Provides additional semantic signal for embedding models.”
↩︎ Giving chunks document-level context“Search children, return parents”
↩︎ Exam trap 1“This approach requires downstream processing to leverage the metadata:”
↩︎ Exam trap 2“Preserve document structure (headings, lists, tables).”
↩︎ Checkpoint“Search children, return parents”
↩︎ Prediction“# Search children, return parents”
↩︎ Checkpoint“For semantic metadata: Use reranking with columns_to_rerank parameter to consider these columns.”
↩︎ Checkpoint - 3.
“Content: the raw chunk text. Same value as the chunk_to_retrieve field for the chunk.”
↩︎ ai_prep_search: context-enriched chunks in SQL“natural-language questions the table is capable of answering, used to improve retrieval recall for table content”
↩︎ ai_prep_search: context-enriched chunks in SQL“Without the schema option, an LLM discovers field names per document”
↩︎ ai_prep_search: context-enriched chunks in SQL“using chunk_to_embed as the embedding column and chunk_id as the primary key”
↩︎ Checkpoint