CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 13/56

    Advanced RAG Chunking: Parent-Child, Contextual Metadata and ai_prep_search

    Design retrieval systems using advanced chunking strategies

    12 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Clean source text so that chunk boundaries follow content rather than page furniture
    • Design parent-child (small-to-big) chunking that searches precise chunks and returns wider context
    • Add semantic context to chunks inline or as metadata columns, and know what each option requires downstream
    • Read the chunk_to_embed and chunk_to_retrieve output of ai_prep_search and know which column feeds the index

    1.Clean the text so chunks carry content, not boilerplate

    A chunking strategy can only be as good as the text it receives. Page headers, footers and page numbers repeat on every page of a PDF. If they reach the chunker, they end up inside chunks, where they add noise to the embedding and can break a sentence across a boundary. The unstructured-pipeline guide tells you to "Preprocess the extracted text to remove irrelevant or noisy information, such as headers, footers, or special characters."

    The AI Search retrieval quality guide pairs cleaning with chunking in three rules: remove boilerplate (headers, footers, page numbers), preserve document structure (headings, lists, tables), and maintain semantic boundaries when chunking. The second rule matters for the advanced designs on this page. Structure-aware chunking and section-header metadata only work if the headings survive cleaning.

    Parsing comes before all of this. For PDFs and complex documents, Databricks recommends ai_parse_document, because poor parsing (missing tables, broken formatting) directly hurts retrieval quality.

    Checkpoint 1 of 6· Check yourself

    Which preprocessing step goes AGAINST the retrieval quality guide's cleaning advice?

    Sources12

    2.Parent-child chunking: search small, return big

    The previous page described a size trade-off: small chunks locate facts precisely, while large chunks keep context. Parent-child chunking uses both instead of compromising between them. You cut each document into large parent chunks, then cut each parent into small child chunks. Each child row stores its own text plus its parent's text. The index embeds and searches the children, and the query asks for the parent column back, so the LLM gets the surrounding context the small match lacked.

    The retrieval quality guide lists this among the advanced approaches that take more effort but can have a bigger impact, alongside semantic chunking:

    Parent-child chunking: 2048-token parents for context, 512-token children for precision, with the parent text returned at query timepython
    # Record child and parent chunks in your source table
    for parent_chunk in create_chunks(doc, size=2048):  # Large for context
        for child_chunk in create_chunks(parent_chunk, size=512):  # Small for precision
            source_table.append({"text": child_chunk, "parent_text": parent_chunk})
    
    # Search children, return parents
    results = index.similarity_search(
        query_text="Is attention all you need?",
        num_results=10,
        columns=["text", "parent_text"]
    )

    Checkpoint 2 of 6· Check yourself

    In the parent-child code sample, what happens at query time?

    Sources2

    3.Giving chunks document-level context

    Once a chunk is cut out of its document, it often loses the context that made it findable. A paragraph about "adjustment procedures" doesn't say which manual or which system it belongs to. The retrieval quality guide recommends adding semantic context because it gives embedding models an extra semantic signal, gives rerankers more to score on, and helps with queries that refer to document-level concepts. It offers two options.

    Option 1: put the context inside the chunk. Prepend the document title, summary and section name to the chunk text before embedding:

    Option 1: prepend document-level context to every chunkpython
    # Prepend document summary to each chunk
    chunk_with_context = f"""
    Document: {doc_title}
    Summary: {doc_summary}
    Section: {section_name}
    {chunk_content}
    """

    Option 2: store the context in separate metadata columns. The chunk text stays as it is, and the summary, section and keywords sit beside it:

    Option 2: semantic metadata stored next to the chunk instead of inside itpython
    # Store semantic metadata for reranker to use
    metadata = {
        "doc_summary": "Technical manual for brake system maintenance",
        "section": "Emergency brake adjustment procedures",
        "keywords": ["brake", "safety", "adjustment"]
    }

    The two options differ in a way that matters. Option 1 changes the embedded text, so the context affects the vector directly. Option 2 doesn't touch the vector, and the guide warns that it "requires downstream processing to leverage the metadata." For semantic metadata, that means reranking with the columns_to_rerank parameter. For keyword-only metadata, it means hybrid (full-text) search against those fields. The unstructured-pipeline guide makes the general point too: store metadata alongside chunks or their embeddings, because it narrows down what gets retrieved.

    Checkpoint 3 of 6· Check yourself

    A team stores doc_summary and section as separate metadata columns next to each chunk, with no other changes. Retrieval doesn't improve. What is missing?

    Checkpoint 4 of 6· Exam question

    During evaluation, a RAG engineer notices that answers referencing information near a chunk boundary are frequently incomplete, because the fact needed to answer the question is split across two adjacent chunks. Which change to the chunking configuration best addresses this specific problem?

    Sources21

    Databricks provides a SQL function that applies Option 1 for you. ai_prep_search takes the output of ai_parse_document and returns chunks. Each chunk comes in two versions: chunk_to_retrieve, the raw chunk text, and chunk_to_embed, the same text wrapped in labelled document context. The documentation states that the Content field of chunk_to_embed is "the raw chunk text. Same value as the chunk_to_retrieve field for the chunk." The template looks like this:

    The chunk_to_embed template: labelled fields, a context sentence and table aids around the raw chunktext
    Document Title: {doc_title}
    Page Header: {page_header}
    Page Footer: {page_footer}
    Section Header: {section_header}
    Caption: {caption}
    Footnote: {footnote}
    Page Number: {page_number}
    {additional_document_fields}
    
    {document_context_sentence}
    
    Table summary: {table_summary}
    
    Content:
    {chunk_to_retrieve}
    
    Related questions:
    {qa_text}

    Three details are worth knowing. The fixed fields (title, header, footer, section header and so on) always appear with their label, left empty when there's no value. For chunks that contain a table, the function adds an LLM-written table summary and related natural-language questions, which are meant to improve recall for table content. Without the schema option, an LLM decides which document fields to extract for each document. Passing schema makes every chunk carry the same keys. The flattened output becomes a Delta table that can feed an AI Search index:

    Columns of the flattened ai_prep_search output that the documentation gives a role for
    ColumnTypeRole
    chunk_idSTRINGPrimary key of the AI Search index
    chunk_to_embedSTRINGEmbedding column: the chunk wrapped in labelled document context
    chunk_to_retrieveSTRINGThe raw chunk text
    metadataVARIANTExtracted document fields, such as the keys passed in the schema option
    source_uriSTRINGPath of the source file read from the volume

    Checkpoint 5 of 6· Fill the gap

    Which function turns the parsed documents into chunks in this pipeline?

    prepped_documents AS (
      SELECT
        path,
         ? (parsed) AS result
      FROM parsed_documents
    )
    SELECT
      chunk.value:chunk_id::STRING AS chunk_id,
      chunk.value:chunk_position::INT AS chunk_position,
      chunk.value:chunk_to_retrieve::STRING AS chunk_to_retrieve,
      chunk.value:chunk_to_embed::STRING AS chunk_to_embed,
      chunk.value:metadata AS metadata,
      prepped_documents.path AS source_uri
    FROM
      prepped_documents,
      LATERAL variant_explode(prepped_documents.result:document.contents) AS chunk;

    Checkpoint 6 of 6· Check yourself

    You're creating an AI Search index from the flattened ai_prep_search table. Which column should be the embedding column?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.In parent-child retrieval, the large parent chunks are embedded and searched because they carry more context.Why is that wrong?

      The small children are what get embedded and searched, for precision. The parent text is stored next to each child and returned as context.

      Covered in Parent-child chunking: search small, return big

    2. 2.Storing a document summary in a separate metadata column improves vector retrieval on its own.Why is that wrong?

      Separate metadata columns don't change the embedded text. They only help when a reranker (columns_to_rerank) or hybrid keyword search is set up to use them.

      Covered in Giving chunks document-level context

    Practise it for real

    Turn documents in a Unity Catalog volume into context-enriched chunk rows with ai_parse_document and ai_prep_search, then compare what gets embedded with what gets retrieved.

    1. 1.Run a query that reads your volume with READ_FILES('/Volumes/mydata/documents/', format => 'binaryFile') and selects path plus ai_parse_document(content) AS parsed, wrapped in a parsed_documents CTE.

      Why: ai_prep_search expects parsed documents, and clean parsing comes before any chunking.

      You should see: One row per file, with a parsed result for each document.

    2. 2.Add a prepped_documents CTE that selects path and ai_prep_search(parsed) AS result from parsed_documents.

      Why: This step does the chunking and builds the context-enriched chunk_to_embed text.

      You should see: One result per document, holding a document.contents array of chunks.

    3. 3.Flatten with LATERAL variant_explode(prepped_documents.result:document.contents) AS chunk, and project chunk_id, chunk_position, chunk_to_retrieve, chunk_to_embed, metadata and path AS source_uri.

      Why: An AI Search source table needs one row per chunk, with a primary key and an embedding column.

      You should see: Rows with the columns chunk_id STRING, chunk_position INT, chunk_to_retrieve STRING, chunk_to_embed STRING, metadata VARIANT, source_uri STRING.

    4. 4.Compare chunk_to_embed and chunk_to_retrieve for a few rows, including a chunk that contains a table.

      Why: Seeing the labelled header fields, context sentence and table aids shows exactly what extra signal the embedding gets.

      You should see: chunk_to_embed starts with labelled fields such as Document Title and Section Header, with the raw chunk_to_retrieve text under Content. Table chunks also carry a table summary and related questions.

    Stuck? Get a nudge

    If the metadata fields vary from document to document, rerun ai_prep_search with map('schema', '["company", "document_type", "fiscal_year"]') so that every chunk carries the same keys.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Preprocess the extracted text to remove irrelevant or noisy information, such as headers, footers, or special characters.”
      ↩︎ Clean the text so chunks carry content, not boilerplate
      “Storing metadata alongside chunked documents or their corresponding embeddings is essential for optimal performance.”
      ↩︎ Giving chunks document-level context
    2. 2.
      “Remove boilerplate (headers, footers, page numbers).”
      ↩︎ Clean the text so chunks carry content, not boilerplate
      “Poor parsing (missing tables, broken formatting) directly impacts retrieval quality.”
      ↩︎ Clean the text so chunks carry content, not boilerplate
      “Parent-child chunking (small-to-big retrieval):”
      ↩︎ Parent-child chunking: search small, return big
      “These techniques require more effort but can have a bigger impact:”
      ↩︎ Parent-child chunking: search small, return big
      “Provides additional semantic signal for embedding models.”
      ↩︎ Giving chunks document-level context
      “Search children, return parents”
      ↩︎ Exam trap 1
      “This approach requires downstream processing to leverage the metadata:”
      ↩︎ Exam trap 2
      “Preserve document structure (headings, lists, tables).”
      ↩︎ Checkpoint
      “Search children, return parents”
      ↩︎ Prediction
      “# Search children, return parents”
      ↩︎ Checkpoint
      “For semantic metadata: Use reranking with columns_to_rerank parameter to consider these columns.”
      ↩︎ Checkpoint
    3. 3.
      “Content: the raw chunk text. Same value as the chunk_to_retrieve field for the chunk.”
      ↩︎ ai_prep_search: context-enriched chunks in SQL
      “natural-language questions the table is capable of answering, used to improve retrieval recall for table content”
      ↩︎ ai_prep_search: context-enriched chunks in SQL
      “Without the schema option, an LLM discovers field names per document”
      ↩︎ ai_prep_search: context-enriched chunks in SQL
      “using chunk_to_embed as the embedding column and chunk_id as the primary key”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 6 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.