CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 2 · Lesson 13/56

    RAG Chunking Strategies: Fixed-Size, Paragraph, Format-Specific and Semantic

    Design retrieval systems using advanced chunking strategies

    10 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Place chunking correctly in a RAG data pipeline and name the five factors that shape a chunking design
    • Choose between fixed-size, paragraph-based, format-specific and semantic chunking for a given document structure
    • Size chunks against embedding-model token limits and the LLM context window, and know which optimizations outrank chunk-size tuning

    Key concept

    Semantically coherent chunk — A chunk is the unit your retriever returns, so it should hold related information and still make sense on its own. Every chunking strategy and size choice is a way of getting closer to that while staying inside model limits.

    1.Where chunking sits and what it decides

    A RAG data pipeline turns raw documents into something a retriever can query. The Databricks unstructured-pipeline guide breaks it into these stages: choosing the corpus, preprocessing (parsing, enrichment, metadata extraction, deduplication, filtering), chunking, embedding, and indexing. Chunking runs once the text is clean and before anything is embedded. This is where you decide the unit of text that retrieval will hand to the LLM.

    Chunks have two jobs. They must be small enough that several of them fit in the LLM's context together, and focused enough that a retrieved chunk isn't padded with off-topic text. The guide puts it as chunking "ensures that retrieved data fits in the LLM's context while minimizing the inclusion of distracting or irrelevant information." Everything downstream sees only what the chunks contain, so the guide calls chunking "one of the first layers of optimization in an RAG application."

    Checkpoint 1 of 5· Put it in order

    Put these data-pipeline stages in the order the Databricks guide describes

    1. 1.Parse the raw documents into a more structured format
    2. 2.Break the cleaned documents into chunks
    3. 3.Convert each chunk into a vector with an embedding model
    4. 4.Remove duplicates and filter out unwanted documents

    The guide names five factors to weigh when you design chunking:

    - Chunking strategy: how the text is divided (by sentence, paragraph, character or token count, or document-specific structure). - Chunk size: small chunks focus on specific details but lose surrounding context. Large chunks keep context but bring in irrelevant material and cost more to compute. - Overlap: repeating some text between adjacent chunks so that information sitting on a boundary isn't lost. - Semantic coherence: chunks that can stand on their own, built by following paragraphs, sections or topic boundaries. - Metadata: source document name, section heading or product names attached to each chunk to help match queries.

    Chunking isn't a self-contained problem. The quality overview points out that retrieval quality depends on choices in both the pipeline and the chain, and the pipeline choices it lists include the chunking strategy.

    Sources12

    2.Four strategies, matched to document structure

    The guide says outright that "There is no one-size-fits-all approach." The right method depends on your use case and on how your documents are structured. It groups the options into four families, from cheapest to most involved.

    Fixed-size chunking splits text every N characters or tokens. It's fast to set up, but the cuts land wherever the count runs out, often mid-sentence or mid-table, so the chunks are rarely coherent. The guide's verdict: "This approach rarely works for production-grade applications."

    Paragraph-based chunking follows the paragraph breaks the author already wrote. Paragraphs usually hold related information, so this preserves coherence at little cost.

    Format-specific chunking uses the structure built into formats like Markdown or HTML. Headers or sections become the chunk boundaries, so a chunk lines up with a section a human would recognise.

    Semantic chunking looks at the content to find where the topic changes, for example with topic modeling or by comparing sentence embeddings. The AI Search retrieval quality guide describes it as grouping sentences by similarity instead of by fixed size, which keeps related ideas together. It takes more effort, but the boundaries match the text's natural divisions.

    Chunking strategies compared, with the LangChain splitter the Databricks guide cites for each
    StrategyHow boundaries are chosenExample splitterTrade-off
    Fixed-sizeA predetermined number of characters or tokensLangChain CharacterTextSplitterQuick and easy, but typically not semantically coherent
    Paragraph-basedNatural paragraph boundaries in the textLangChain RecursiveCharacterTextSplitterHelps preserve coherence, since paragraphs often hold related information
    Format-specificInherent structure such as Markdown headers or HTML sectionsLangChain MarkdownHeaderTextSplitterOnly available where the format carries structure
    SemanticTopic shifts in the contentLangChain SemanticChunkerMore involved, but aligned with natural semantic divisions

    Checkpoint 2 of 5· Match them up

    Match each chunking strategy to how it sets chunk boundaries

    Tap a term, then the definition that fits it.

    Whichever strategy you pick, overlap is a separate setting. Repeating a little text between neighbouring chunks protects facts that would otherwise be cut in two. According to the guide, overlap can preserve continuity and context and improve retrieval results.

    Checkpoint 3 of 5· Exam question

    A team is building a RAG chatbot over internal engineering documentation written in Markdown, where every page uses consistent H1/H2/H3 headers to organize topics such as installation, configuration, and troubleshooting. Which chunking strategy is best suited to this document set to preserve topical boundaries?

    Sources13

    3.Sizing chunks against embedding and LLM limits

    Two limits cap chunk size. The first is the embedding model's maximum token count. The guide's example is bge-large-en-v1.5, which has a limit of 512 tokens, and chunks longer than that get truncated. The second is the LLM context window. The retrieval quality guide notes that retrieved chunks must fit inside it, and you usually retrieve several at a time, so the space is shared.

    A large embedding limit doesn't mean your chunks should be that large. The AI Search example notebooks use a model that accepts 8,192 tokens, yet they still advise splitting into smaller chunks so the reasoning model gets a wider variety of examples. The OpenAI-embedding example splits on a fixed token count of 1,024:

    Token-based fixed-size splitting from the AI Search external-embedding example (max_chunk_tokens is set to 1024 above this function)python
    def chunk_text(text):
        # Encode and then decode within the UDF
        tokens = encoding.encode(text)
        chunks = []
        while tokens:
            chunk_tokens = tokens[:max_chunk_tokens]
            chunk_text = encoding.decode(chunk_tokens)
            chunks.append(chunk_text)
            tokens = tokens[max_chunk_tokens:]
        return chunks

    For experiments, the retrieval quality guide suggests testing 256 tokens for precise fact retrieval, 512 as a middle ground, and 1,024 for more context per chunk. Smaller chunks pinpoint information better but can lose context. Larger chunks keep context but make the relevant passage harder to locate.

    The guide also warns against spending too long on size. It calls chunk size an active research area, then says: "Instead of over-optimizing chunk size, focus on:" extracting metadata (entities, topics, categories) for filtering, parsing documents cleanly, and adding semantic metadata such as summaries and section headers to chunks.

    Checkpoint 4 of 5· Check yourself

    A team has spent two sprints trying 256, 512 and 1,024-token chunks, and the gains are small. What does the AI Search retrieval quality guide say to focus on instead?

    Checkpoint 5 of 5· Exam question

    A support knowledge base consists of long compiled FAQ articles where a single page can cover several unrelated subtopics with no headers or paragraph markers to indicate where one topic ends and another begins. Which chunking approach directly addresses this document characteristic?

    Sources143

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Fixed-size character or token splitting is a sound production default because it's simple and predictable.Why is that wrong?

      Arbitrary character or token counts usually produce chunks that aren't semantically coherent. The guide says this approach rarely works in production; prefer paragraph, format-specific or semantic boundaries.

      Covered in Four strategies, matched to document structure

    2. 2.Bigger chunks are always safer because they keep more context.Why is that wrong?

      Chunks over the embedding model's token limit are truncated, and retrieved chunks must also fit in the LLM context window together. Bigger chunks keep context but make relevant information harder to pinpoint.

      Covered in Sizing chunks against embedding and LLM limits

    3. 3.Tuning chunk size is the main lever for improving retrieval quality.Why is that wrong?

      The retrieval quality guide calls chunk size an active research area and points to metadata extraction, high-quality parsing and semantic metadata as more impactful.

      Covered in Sizing chunks against embedding and LLM limits

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “ensures that retrieved data fits in the LLM's context while minimizing the inclusion of distracting or irrelevant information”
      ↩︎ Where chunking sits and what it decides
      “making it one of the first layers of optimization in an RAG application”
      ↩︎ Where chunking sits and what it decides
      “There is no one-size-fits-all approach.”
      ↩︎ Four strategies, matched to document structure
      “Formats such as Markdown or HTML have an inherent structure that can define chunk boundaries”
      ↩︎ Four strategies, matched to document structure
      “Overlapping can ensure continuity and context preservation across chunks and improve the retrieval results.”
      ↩︎ Four strategies, matched to document structure
      “For example, bge-large-en-v1.5 has a maximum token limit of 512.”
      ↩︎ Sizing chunks against embedding and LLM limits
      “aim to create semantically coherent chunks that contain related information but can stand independently as a meaningful unit of text”
      ↩︎ Key concept
      “This approach rarely works for production-grade applications.”
      ↩︎ Exam trap 1
      “removing duplicates, and filtering out unwanted information, the next step is to break it down into smaller, manageable units called chunks”
      ↩︎ Checkpoint
      “Use the natural paragraph boundaries in the text to define chunks.”
      ↩︎ Checkpoint
      “If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
      ↩︎ Prediction
    2. 2.
      “Retrieval quality can be influenced by both the data pipeline (for example, parsing/chunking strategy, metadata strategy, embedding model)”
      ↩︎ Where chunking sits and what it decides
    3. 3.
      “Use embeddings to find natural semantic boundaries.”
      ↩︎ Four strategies, matched to document structure
      “Smaller chunks: Better localization of specific information, but may lose context.”
      ↩︎ Sizing chunks against embedding and LLM limits
      “Context limits: Must fit within LLM context window when retrieving multiple chunks.”
      ↩︎ Exam trap 2
      “Instead of over-optimizing chunk size, focus on:”
      ↩︎ Exam trap 3
      “Information extraction for metadata: Extract entities, topics, and categories to enable precise filtering.”
      ↩︎ Checkpoint
    4. 4.
      “split the data into smaller context chunks so that you can feed a wider variety of examples into the reasoning model”
      ↩︎ Sizing chunks against embedding and LLM limits

    Continue to page 2 of 2

    Advanced RAG Chunking: Parent-Child, Contextual Metadata and ai_prep_search

    Spotted a mistake, or was something unclear? Tell us.