CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 7/56

    Chunking Strategies for Document Structure and Model Limits

    Apply a chunking strategy for a given document structure and model constraints

    9 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Place chunking correctly in a RAG data pipeline and name the factors that shape a chunking decision
    • Choose between fixed-size, paragraph-based, format-specific and semantic chunking based on how a document is structured
    • Size chunks against the embedding model's token limit and the LLM's context window

    Key concept

    Chunk as a constrained retrieval unit — A chunk is the piece of a document that gets embedded and retrieved. It has to fit within the embedding model's token limit and, along with the other retrieved chunks, within the LLM's context. It should also make sense on its own, so the boundaries need to follow the document's real structure.

    1.Where chunking sits, and what you are deciding

    A RAG application over unstructured content needs a data pipeline. The pipeline parses raw files, cleans them, cuts them into chunks, embeds each chunk and indexes the vectors. Chunking comes after the text has been extracted and cleaned, and before anything is embedded. Because of where it sits, every later stage works on whatever chunks you produce. The Databricks guidance calls chunking one of the first layers of optimization in a RAG application, because it directly controls what the LLM is given at answer time.

    Checkpoint 1 of 6· Put it in order

    Put these data pipeline stages in the order the Databricks guidance describes them

    1. 1.Break the cleaned text into chunks
    2. 2.Parse the raw data into a more structured format
    3. 3.Remove duplicate documents
    4. 4.Filter out unwanted information

    "Choose a chunk size" is only one part of a chunking decision. The guidance lists five factors. Strategy is the method used to cut the text, for example by sentences, paragraphs, character or token counts, or document-specific rules. Chunk size is a trade-off: smaller chunks focus on specific details but lose surrounding context, and larger chunks keep context but can pull in irrelevant text and cost more to process. Overlap between neighbouring chunks keeps information from being lost at a boundary. Semantic coherence means each chunk should hold related information and still make sense on its own. Metadata, such as the source document name, section heading or product name, gives retrieval queries more to match against.

    Sources1

    2.Matching the strategy to the document's structure

    The Databricks guidance does not name a single best method. It says the right chunk size and method depend on the use case and on the data. It groups the options into four families, and each one ties a kind of document structure to a kind of boundary. Fixed-size splitting ignores structure entirely: it is quick to set up, but it usually does not produce coherent chunks. Paragraph-based splitting uses the paragraph breaks the author already wrote. Format-specific splitting uses explicit markup such as headers. Semantic chunking looks at the content to find where the topic changes, which helps with prose that has little visible structure.

    The four chunking families, what defines a boundary, and the LangChain example the guidance gives for each
    StrategyBoundary comes fromExample toolTrade-off
    Fixed-sizeA set number of characters or tokensCharacterTextSplitterQuick and easy to set up, but chunks are typically not semantically coherent
    Paragraph-basedNatural paragraph boundariesRecursiveCharacterTextSplitterHelps keep related information together
    Format-specificInherent markup such as Markdown headers or HTML sectionsMarkdownHeaderTextSplitter, HTML header/section splittersOnly works where the format carries structure
    SemanticTopic shifts found by analysing the contentSemanticChunkerMore work, but chunks follow natural semantic divisions

    Checkpoint 2 of 6· Match them up

    Match each chunking strategy to the LangChain tool the Databricks guidance gives as its example

    Tap a term, then the definition that fits it.

    Checkpoint 3 of 6· Exam question

    A data engineering team ingests thousands of unstructured plain-text support-ticket transcripts and applies a chunking strategy that splits every document into fixed 512-character blocks with no attention to sentence or paragraph boundaries. During evaluation, retrieved chunks frequently cut sentences in half, and the RAG application returns incomplete or confusing context to the LLM. What should the team do to address this?

    Sources1

    3.Sizing chunks against the embedding model and the LLM

    Two model limits cap chunk size. The first is the embedding model's maximum token limit. The guidance gives bge-large-en-v1.5 as an example, with a limit of 512 tokens. Any chunk longer than the limit is cut off before it is embedded. The second is the LLM's context window. You usually retrieve several chunks for each query, and all of them have to fit into the prompt together. So the limit that matters is chunk size multiplied by the number of chunks you retrieve, not the size of one chunk.

    Fitting under the embedding model's limit is necessary but not enough. The AI Search embedding examples point out that the GTE model, and the OpenAI model in the external example, support up to 8192 tokens. Databricks still recommends splitting data into smaller chunks, so that the model generating the answer receives a wider range of examples. The OpenAI example caps chunks at 1024 tokens.

    Checkpoint 4 of 6· Fill the gap

    This chunking function comes from the external-embedding-model example. The model supports 8192 tokens, but the example follows Databricks' advice to use smaller chunks. Which value fills the blank?

    max_chunk_tokens =  ? 
    encoding = tiktoken.get_encoding("cl100k_base")
    
    
    def chunk_text(text):
        # Encode and then decode within the UDF
        tokens = encoding.encode(text)
        chunks = []
        while tokens:
            chunk_tokens = tokens[:max_chunk_tokens]
            chunk_text = encoding.decode(chunk_tokens)
            chunks.append(chunk_text)
            tokens = tokens[max_chunk_tokens:]
        return chunks

    For choosing a size within those limits, the AI Search retrieval quality guide suggests three starting points to test. It also says chunk size optimization is still an active area of research, so treat these as experiments rather than answers.

    Starting chunk sizes the AI Search retrieval quality guide suggests testingpython
    # Common configurations to test
    small_chunks = 256   # Better for precise fact retrieval
    medium_chunks = 512  # Balanced approach
    large_chunks = 1024  # More context per chunk

    Checkpoint 5 of 6· Check yourself

    A RAG chain retrieves ten chunks for each question. You are thinking of moving from 512-token to 1024-token chunks. Which constraint does the guidance say you must check?

    Checkpoint 6 of 6· Exam question

    An engineer builds a RAG pipeline over long onboarding guides. During testing, they notice that a key instruction spanning the end of one chunk and the beginning of the next is never retrieved in full: the retriever returns one chunk or the other, but neither chunk alone contains the complete instruction. Which adjustment to the chunking configuration most directly addresses this failure?

    Sources12

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Fixed-size chunking by character or token count is a sound default for a production RAG system.Why is that wrong?

      Fixed-size splitting is quick to set up, but it typically does not produce semantically coherent chunks. The guidance says it rarely works in production. Use the document's paragraphs, markup or topics to set boundaries instead.

      Covered in Matching the strategy to the document's structure

    2. 2.If the embedding model accepts 8192 tokens, chunks should be close to 8192 tokens.Why is that wrong?

      The limit is a ceiling, not a target. Databricks recommends smaller chunks so the reasoning model receives a wider range of examples.

      Covered in Sizing chunks against the embedding model and the LLM

    3. 3.A chunk longer than the embedding model's limit causes an error, so you will notice it.Why is that wrong?

      Oversized chunks are truncated without any error, so the text past the limit never reaches the embedding.

      Covered in Sizing chunks against the embedding model and the LLM

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Smaller chunks may focus on specific details but lose some surrounding contextual information.”
      ↩︎ Where chunking sits, and what you are deciding
      “aim to create semantically coherent chunks that contain related information but can stand independently as a meaningful unit of text”
      ↩︎ Where chunking sits, and what you are deciding
      “Finding the proper chunking method is both iterative and context-dependent. There is no one-size-fits-all approach.”
      ↩︎ Matching the strategy to the document's structure
      “Use the natural paragraph boundaries in the text to define chunks.”
      ↩︎ Matching the strategy to the document's structure
      “For example, bge-large-en-v1.5 has a maximum token limit of 512.”
      ↩︎ Sizing chunks against the embedding model and the LLM
      “Segmenting large documents into smaller, semantically concentrated chunks ensures that retrieved data fits in the LLM's context”
      ↩︎ Key concept
      “This approach rarely works for production-grade applications.”
      ↩︎ Exam trap 1
      “If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
      ↩︎ Exam trap 3
      “After parsing the raw data into a more structured format, removing duplicates, and filtering out unwanted information, the next step is to break”
      ↩︎ Checkpoint
      “Formats such as Markdown or HTML have an inherent structure that can define chunk boundaries”
      ↩︎ Prediction
      “Tools like LangChain's MarkdownHeaderTextSplitter or HTML header/section-based splitters can be used for this purpose.”
      ↩︎ Checkpoint
      “If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
      ↩︎ Prediction
    2. 2.
      “Chunk size optimization remains an active area of research.”
      ↩︎ Sizing chunks against the embedding model and the LLM
      “Context limits: Must fit within LLM context window when retrieving multiple chunks.”
      ↩︎ Checkpoint

    Also cited

    Continue to page 2 of 2

    Advanced Chunking: Semantic, Parent-Child and Context-Enriched Chunks

    Spotted a mistake, or was something unclear? Tell us.