CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 17/56

    RAG Chunking Strategies: Size, Overlap, Structure and Model Limits

    Select chunking strategy based on model & retrieval evaluation

    9 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Name the four broad families of chunking strategy and the LangChain splitter that represents each
    • Explain the trade-offs of chunk size and overlap, and how semantic coherence and metadata affect retrieval
    • Use the embedding model's token limit and the LLM's context window to bound chunk size
    • Explain when parent-child chunking and adding semantic context to chunks are worth trying

    Key concept

    Chunking as a tunable retrieval knob — No single chunking strategy is correct for every RAG application. Strategy, size, overlap and metadata are settings you pick for your data and then refine by testing, because they decide which text the retriever can return to the LLM.

    1.Why chunking is an early quality lever

    In a RAG data pipeline, documents are parsed, deduplicated and filtered, then split into chunks. Each chunk is embedded and indexed. The aim is chunks that are small and focused enough to fit in the LLM's context without bringing in distracting text. The retriever can only return chunks, so the way you split documents limits what the LLM ever sees. Databricks calls chunking one of the first layers of optimization in a RAG application.

    The cookbook lists five factors to decide when you chunk:

    - Strategy: how the text is divided, from sentence, paragraph or token-count splits to document-specific methods. - Size: small chunks pinpoint details but drop the surrounding context. Large chunks keep context but bring in noise. - Overlap: sharing some text between neighbouring chunks keeps continuity, so a fact that falls on a boundary isn't lost. - Semantic coherence: each chunk should hold related information and still make sense on its own. - Metadata: the source document name, section heading or product names help match queries to chunks.

    Changing any one of these changes the result set the LLM receives.

    Checkpoint 1 of 5· Check yourself

    Facts that span two adjacent chunks are often missing from retrieved results. Which chunking factor is designed for this problem?

    Sources1

    2.The four families of chunking strategy

    Databricks groups chunking strategies into four families. Moving down the list, each one splits text along more meaningful boundaries and takes more effort to set up. Fixed-size splitting is quick to configure, but it cuts text at arbitrary points, so chunks are often not semantically coherent. Paragraph-based splitting uses the document's own paragraph breaks. Format-specific splitting uses structure such as Markdown or HTML headers. Semantic chunking analyses the content itself, for example with topic modeling, to find where the topic changes.

    Chunking strategy families and the example LangChain components the Databricks cookbook names
    StrategyHow boundaries are chosenExample componentMain caveat
    Fixed-sizeA set number of characters or tokensCharacterTextSplitterRarely works for production-grade applications
    Paragraph-basedNatural paragraph boundariesRecursiveCharacterTextSplitterDepends on the document having meaningful paragraphs
    Format-specificInherent structure such as Markdown or HTML headersMarkdownHeaderTextSplitterOnly works for formats that have that structure
    SemanticTopic shifts found by analysing the contentSemanticChunkerMore involved to set up than basic approaches

    Checkpoint 2 of 5· Match them up

    Match each document situation to the chunking strategy family the docs associate with it

    Tap a term, then the definition that fits it.

    Checkpoint 3 of 5· Exam question

    A team building a RAG assistant over a corpus of legal contracts is currently using fixed-size, token-count chunking (500 tokens, no overlap). Retrieval evaluation shows that many chunks split individual contract clauses across two consecutive chunks, so the retriever often returns only half of a relevant clause. The contracts have clear, consistent section and clause headers (e.g., "4.2 Termination"). Which change to the chunking strategy is most likely to fix this specific issue?

    Sources1

    3.Model limits that bound chunk size

    Two models put limits on your chunk size.

    The embedding model. Every embedding model has a maximum number of input tokens. If a chunk is longer, the model truncates it, and the text past the limit never reaches the vector. For example, bge-large-en-v1.5 accepts at most 512 tokens. A 1,000-token chunk would be embedded from only its first part, even though all of the text is stored.

    The generation LLM. The retriever usually returns several chunks, and all of them have to fit in the LLM's context window together with the prompt. So the right chunk size also depends on how many chunks you plan to retrieve. The embedding model's token limit is a hard maximum. It does not tell you the best chunk size.

    Checkpoint 4 of 5· Check yourself

    A pipeline embeds 900-token chunks with bge-large-en-v1.5. Retrieval misses facts that appear near the end of long chunks. What is the most likely cause?

    Sources123

    4.Beyond chunk size: semantic context and parent-child chunks

    The AI Search retrieval quality guide calls chunk size optimization an active area of research. It recommends that you not spend most of your effort on chunk size. Higher-impact work includes extracting metadata such as entities and topics for filtering, parsing documents well (for example with ai_parse_document), and adding semantic context to chunks. One way to add context is to put the document title, a summary and the section name at the start of each chunk, so the embedding has signal from the whole document.

    Adding document-level context to the start of each chunk before embeddingpython
    # Prepend document summary to each chunk
    chunk_with_context = f"""
    Document: {doc_title}
    Summary: {doc_summary}
    Section: {section_name}
    {chunk_content}
    """

    The guide lists two advanced strategies that take more effort but can have a bigger impact. Semantic chunking uses embeddings to group sentences by similarity, so related ideas stay together. Parent-child chunking, also called small-to-big retrieval, sidesteps the trade-off between small and large chunks. You search small child chunks, which locate information precisely, and return the larger parent chunk, which gives the LLM context.

    Checkpoint 5 of 5· Fill the gap

    In this parent-child retrieval sample, which column must be requested so the LLM receives the large context chunk?

    # Record child and parent chunks in your source table
    for parent_chunk in create_chunks(doc, size=2048):  # Large for context
        for child_chunk in create_chunks(parent_chunk, size=512):  # Small for precision
            source_table.append({"text": child_chunk, "parent_text": parent_chunk})
    
    # Search children, return parents
    results = index.similarity_search(
        query_text="Is attention all you need?",
        num_results=10,
        columns=["text", " ? "]
    )

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Fixed-size token splitting is an acceptable default for a production RAG application.Why is that wrong?

      Fixed-size splitting is quick to set up, but it seldom produces semantically coherent chunks. The cookbook says it rarely works in production.

      Covered in The four families of chunking strategy

    2. 2.If the embedding model supports 8192 tokens, chunks should be 8192 tokens long.Why is that wrong?

      The token limit is a hard maximum, not a target. Databricks still recommends smaller chunks so the LLM receives a wider variety of examples.

      Covered in Model limits that bound chunk size

    3. 3.Tuning chunk size is the highest-impact way to improve retrieval.Why is that wrong?

      The retrieval quality guide recommends spending effort on metadata extraction, high-quality parsing and adding semantic context to chunks instead of over-tuning chunk size.

      Covered in Beyond chunk size: semantic context and parent-child chunks

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The choices made on chunking will directly affect the retrieved data the LLM provides”
      ↩︎ Why chunking is an early quality lever
      “Smaller chunks may focus on specific details but lose some surrounding contextual information.”
      ↩︎ Why chunking is an early quality lever
      “Relevant metadata, such as the source document name, section heading, or product names, can improve retrieval.”
      ↩︎ Why chunking is an early quality lever
      “Use the natural paragraph boundaries in the text to define chunks.”
      ↩︎ The four families of chunking strategy
      “determine the most appropriate chunk boundaries based on topic shifts”
      ↩︎ The four families of chunking strategy
      “For example, bge-large-en-v1.5 has a maximum token limit of 512.”
      ↩︎ Model limits that bound chunk size
      “Finding the proper chunking method is both iterative and context-dependent. There is no one-size-fits-all approach.”
      ↩︎ Key concept
      “This approach rarely works for production-grade applications.”
      ↩︎ Exam trap 1
      “Larger chunks may capture more context but can include irrelevant information or be computationally expensive.”
      ↩︎ Prediction
      “Overlapping can ensure continuity and context preservation across chunks and improve the retrieval results.”
      ↩︎ Checkpoint
      “Formats such as Markdown or HTML have an inherent structure that can define chunk boundaries (for example, markdown headers).”
      ↩︎ Checkpoint
      “If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
      ↩︎ Checkpoint
    2. 2.
      “Context limits: Must fit within LLM context window when retrieving multiple chunks.”
      ↩︎ Model limits that bound chunk size
      “Chunk size optimization remains an active area of research.”
      ↩︎ Beyond chunk size: semantic context and parent-child chunks
      “Semantic chunking: Group sentences by similarity rather than fixed size.”
      ↩︎ Beyond chunk size: semantic context and parent-child chunks
      “Instead of over-optimizing chunk size, focus on:”
      ↩︎ Exam trap 3
    3. 3.
      “The GTE model supports up to 8192 tokens. However, Databricks recommends that you split the data into smaller context chunks”
      ↩︎ Model limits that bound chunk size
      “The GTE model supports up to 8192 tokens. However, Databricks recommends that you split the data into smaller context chunks”
      ↩︎ Exam trap 2

    Continue to page 2 of 2

    Selecting a Chunking Strategy from Retrieval Evaluation Metrics

    Spotted a mistake, or was something unclear? Tell us.