CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 7/56

    Advanced Chunking: Semantic, Parent-Child and Context-Enriched Chunks

    Apply a chunking strategy for a given document structure and model constraints

    9 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why Databricks advises against over-optimizing chunk size, and what to improve instead
    • Design parent-child (small-to-big) chunking that searches small chunks and returns larger ones
    • Use ai_parse_document structure and ai_prep_search to build chunks that carry document context

    1.Beyond chunk size: semantic boundaries

    The AI Search retrieval quality guide pushes back on tuning chunk size alone. It notes that recent research shows embeddings can fail to capture basic information in long contexts, which makes size a subtle choice rather than a dial to max out. It then names areas that usually matter more: extracting entities, topics and categories as metadata for filtering; high-quality parsing, for which it recommends ai_parse_document; and adding semantic metadata such as document summaries and section headers to chunks. Poor parsing, such as missing tables or broken formatting, directly lowers retrieval quality, and no chunk size fixes that.

    Its first advanced technique is semantic chunking. Instead of cutting at a fixed size, it groups sentences by similarity and uses embeddings to find natural semantic boundaries. The result keeps related ideas together and preserves context better. It needs more work than a splitter with a size setting, and the guide puts it among techniques that can have a bigger impact.

    Checkpoint 1 of 5· Check yourself

    How does semantic chunking, as the AI Search retrieval quality guide describes it, decide where one chunk ends?

    Sources1

    2.Parent-child chunking: search small, return big

    The size trade-off from basic chunking still applies: small chunks pinpoint facts, and large chunks keep context. Parent-child chunking, also called small-to-big retrieval, avoids having to choose. Each document is cut into large parent chunks, and each parent is cut again into small child chunks. Both are stored on the same row of the source table. Similarity search runs against the precise child text, and the query also returns the parent text, so the LLM gets the surrounding context.

    Parent-child chunking from the AI Search retrieval quality guide: 2048-token parents, 512-token children, search on children, return parentspython
    # Record child and parent chunks in your source table
    for parent_chunk in create_chunks(doc, size=2048):  # Large for context
        for child_chunk in create_chunks(parent_chunk, size=512):  # Small for precision
            source_table.append({"text": child_chunk, "parent_text": parent_chunk})
    
    # Search children, return parents
    results = index.similarity_search(
        query_text="Is attention all you need?",
        num_results=10,
        columns=["text", "parent_text"]
    )

    Checkpoint 2 of 5· Check yourself

    A team needs precise matching on specific clauses in long contracts, but the LLM also needs the surrounding section to answer correctly. Which design fits the guidance?

    Checkpoint 3 of 5· Exam question

    A team is preparing a corpus of Markdown product documentation, where each file uses `#`, `##`, and `###` headers to organize content into sections and subsections. They want chunks that align with the document's logical structure rather than arbitrary character counts, so that each chunk corresponds to a coherent section. Which chunking approach best fits this requirement?

    Sources1

    3.Structure-aware chunks that carry their own context

    A chunk taken out of a long report can lose information that only appeared elsewhere, such as which company, which section or which year it belongs to. Structure-aware chunking needs that structure to be detected first. ai_parse_document breaks a document into typed elements, including section_header, page_header, page_footer and page_number. Those types let you split at section boundaries and handle repeated page furniture separately from body text. The retrieval quality guide's data-cleaning advice says to remove boilerplate such as headers, footers and page numbers, preserve document structure, and keep semantic boundaries intact when chunking.

    The guide's simplest form of enrichment adds document-level context to the start of each chunk before embedding. This gives the embedding model extra semantic signal and helps with queries about document-level ideas.

    Adding document-level context to the start of each chunk before embeddingpython
    # Prepend document summary to each chunk
    chunk_with_context = f"""
    Document: {doc_title}
    Summary: {doc_summary}
    Section: {section_name}
    {chunk_content}
    """

    Databricks also has a built-in function for this, ai_prep_search (Beta, Databricks Runtime 18.2 or above). It takes the structured output of ai_parse_document, or plain text or markdown. It splits the content into semantic chunks and enriches each chunk with document-level context. Each output chunk has two text fields: chunk_to_retrieve, the raw chunk text, and chunk_to_embed, an enriched string with the document title, page header and footer, section header, caption, footnote and page number. It can also include LLM-discovered or schema-defined document fields, a one-sentence summary of the document and, for tables, a summary plus related questions. When the flattened output is written to a Delta table for an AI Search index, chunk_to_embed is the embedding column and chunk_id is the primary key.

    Checkpoint 4 of 5· Fill the gap

    This example produces search-ready chunks from raw files in a Unity Catalog volume. Which function fills the blank?

    WITH parsed_documents AS (
      SELECT ai_parse_document(content) AS parsed
      FROM READ_FILES('/Volumes/mydata/documents/', format => 'binaryFile')
    )
    SELECT  ? (parsed) AS result
    FROM parsed_documents;

    Checkpoint 5 of 5· Check yourself

    In ai_prep_search output, which field contains the raw chunk text with no added document context?

    Sources213

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Finding the best chunk size is the most effective way to improve RAG retrieval.Why is that wrong?

      The guide warns against over-optimizing chunk size. It points instead to metadata extraction, high-quality parsing, and adding summaries and section headers to chunks.

      Covered in Beyond chunk size: semantic boundaries

    2. 2.ai_prep_search's chunk_to_embed field is just the raw chunk text, so it does not matter which field you index.Why is that wrong?

      chunk_to_embed adds document-level context to the raw text. That is why it is the embedding column, with chunk_id as the primary key.

      Covered in Structure-aware chunks that carry their own context

    Practise it for real

    Turn raw documents in a Unity Catalog volume into one row per context-enriched chunk, ready to use as an AI Search source table.

    1. 1.In a notebook or the SQL editor on Databricks Runtime 18.2 or above, run ai_parse_document over READ_FILES('/Volumes/mydata/documents/', format => 'binaryFile') in a CTE named parsed_documents, replacing the path with your own volume.

      Why: ai_prep_search works best on structured parser output, which marks section headers, page headers and tables.

      You should see: One row per file, with a VARIANT column called parsed.

    2. 2.Wrap the parsed column in ai_prep_search with map('schema', '["company", "document_type", "fiscal_year"]') and alias the result as result.

      Why: The schema option sets which document-level fields are extracted and attached to every chunk.

      You should see: One row per document, with result:document.contents holding an array of chunks.

    3. 3.Use LATERAL variant_explode(prepped_documents.result:document.contents) AS chunk to flatten the array, and select chunk.value:chunk_id::STRING and chunk.value:chunk_to_embed::STRING along with the metadata keys.

      Why: A vector index source table needs one row per chunk, with a primary key and a text column to embed.

      You should see: One row per chunk. chunk_to_embed starts with 'Document Title:' followed by the section header and page-number lines.

    4. 4.Compare chunk_to_embed with chunk_to_retrieve for a chunk that contains a table.

      Why: This shows the enrichment that only appears for tables.

      You should see: chunk_to_embed includes a 'Table summary:' line and a 'Related questions:' block that chunk_to_retrieve does not have.

    Stuck? Get a nudge

    If the call fails on serverless compute, check that the environment version is 3 or above, since this function needs the VARIANT type.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Poor parsing (missing tables, broken formatting) directly impacts retrieval quality.”
      ↩︎ Beyond chunk size: semantic boundaries
      “Semantic metadata: Add document summaries and section headers to chunks.”
      ↩︎ Beyond chunk size: semantic boundaries
      “Parent-child chunking (small-to-big retrieval):”
      ↩︎ Parent-child chunking: search small, return big
      “Remove boilerplate (headers, footers, page numbers).”
      ↩︎ Structure-aware chunks that carry their own context
      “Provides additional semantic signal for embedding models.”
      ↩︎ Structure-aware chunks that carry their own context
      “Instead of over-optimizing chunk size, focus on:”
      ↩︎ Exam trap 1
      “Semantic chunking: Group sentences by similarity rather than fixed size.”
      ↩︎ Checkpoint
      “# Search children, return parents”
      ↩︎ Checkpoint
    2. 2.
      “section_header: A heading or subheading that denotes the start of a section.”
      ↩︎ Structure-aware chunks that carry their own context
    3. 3.
      “the function splits content into semantic chunks, enriches each chunk with document-level context”
      ↩︎ Structure-aware chunks that carry their own context
      “using chunk_to_embed as the embedding column and chunk_id as the primary key”
      ↩︎ Exam trap 2
      “Content: the raw chunk text. Same value as the chunk_to_retrieve field for the chunk.”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.