CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 22/56

    Chunk Size vs Embedding Context Length: Optimization Strategy

    Select an embedding model context length based on source documents, expected queries, and optimization strategy

    7 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why Databricks recommends chunks smaller than an embedding model's maximum context length
    • Choose starting chunk sizes to test and state the trade-off each one makes
    • Use parent-child chunking and semantic metadata to get precision and context without relying on one long embedding window

    1.Why the maximum is a ceiling, not a target

    An embedding model's context length tells you the largest chunk you can send without truncation. It does not tell you the best chunk size. Both Databricks AI Search example notebooks make this point in the same words. The aim is to split data into smaller chunks "so that you can feed a wider variety of examples into the reasoning model for your RAG application."

    The reason is that retrieved chunks are pasted into the LLM's prompt, usually several at a time. The retrieval quality guide includes this as a hard constraint: "Context limits: Must fit within LLM context window when retrieving multiple chunks." The data-pipeline guide describes chunking as a way to ensure that retrieved data fits in the LLM's context while minimizing the inclusion of distracting or irrelevant information. A few huge chunks fill the generation model's context quickly and carry extra text that doesn't help with the question.

    The OpenAI example notebook shows this in its code. The external OpenAI embedding model it uses supports up to 8192 tokens, but the notebook sets its chunk budget well below that limit.

    Checkpoint 1 of 3· Fill the gap

    The OpenAI embedding model in this notebook supports up to 8192 tokens, yet the notebook follows the recommendation to chunk smaller. Which value does it use?

    max_chunk_tokens =  ? 
    encoding = tiktoken.get_encoding("cl100k_base")

    Sources123

    2.Choosing chunk sizes to test

    If the window only sets a ceiling, the actual chunk size has to come from experiments. The data-pipeline guide says the optimal size "depends on the specific use case and the nature of the data being processed". The retrieval quality guide offers three common configurations to test first.

    Starting chunk sizes (in tokens) suggested for experimentationpython
    # Common configurations to test
    small_chunks = 256   # Better for precise fact retrieval
    medium_chunks = 512  # Balanced approach
    large_chunks = 1024  # More context per chunk

    Compare these with the windows from the model list. All three fit within GTE's 8192 tokens. Only 256 and 512 fit within BGE's 512-token window, so a 1024-token experiment on BGE would test truncated chunks rather than the chunk size you meant to test. The candidate model limits which sizes you can test fairly.

    Each size has a cost. The data-pipeline guide says smaller chunks may focus on specific details but lose surrounding context, while "Larger chunks may capture more context but can include irrelevant information or be computationally expensive." A long window does not guarantee good long-chunk embeddings either. The retrieval quality guide notes that "Recent work from DeepMind (LIMIT) shows embeddings can fail to capture basic information in long contexts, making this a nuanced decision." Model size adds another trade-off: "Larger embedding models generally perform better but require more computational resources."

    Checkpoint 2 of 3· Match them up

    Match each chunk configuration to the trade-off it represents

    Tap a term, then the definition that fits it.

    Sources2

    3.Getting precision and context without one long window

    The retrieval quality guide warns against tuning only this one setting: "More impactful optimizations: Instead of over-optimizing chunk size, focus on:". It then lists metadata extraction, high-quality parsing, and semantic metadata. Two of its techniques are especially relevant to context length because they separate what gets embedded from what gets returned to the LLM.

    The first is parent-child chunking (small-to-big retrieval). You embed and search small child chunks for precision, then return the larger parent chunk to give the LLM context.

    Parent-child chunking: 512-token children are embedded and searched, 2048-token parents are returnedpython
    # Record child and parent chunks in your source table
    for parent_chunk in create_chunks(doc, size=2048):  # Large for context
        for child_chunk in create_chunks(parent_chunk, size=512):  # Small for precision
            source_table.append({"text": child_chunk, "parent_text": parent_chunk})

    The second is adding semantic context: putting a document title, summary, and section name in front of each chunk. According to the guide, this "Provides additional semantic signal for embedding models." The guide also describes semantic chunking, which uses embeddings to find natural boundaries so related ideas stay together instead of being cut at an arbitrary token count.

    Checkpoint 3 of 3· Check yourself

    A team keeps changing chunk size and sees little improvement in retrieval. According to the retrieval quality guide, what should they focus on instead?

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The best chunk size equals the embedding model's maximum context length, such as 8192 tokens for GTE.Why is that wrong?

      The maximum is only a ceiling. Databricks recommends smaller chunks so that a wider variety of examples can go to the LLM, and the retrieved chunks must also fit in the LLM's context window.

      Covered in Why the maximum is a ceiling, not a target

    2. 2.A model with a long embedding window will capture everything in a long chunk equally well.Why is that wrong?

      Research cited by Databricks shows embeddings can fail to capture basic information in long contexts, so long chunks are not automatically represented well.

      Covered in Choosing chunk sizes to test

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Context limits: Must fit within LLM context window when retrieving multiple chunks.”
      ↩︎ Why the maximum is a ceiling, not a target
      “Provides additional semantic signal for embedding models.”
      ↩︎ Getting precision and context without one long window
      “More impactful optimizations: Instead of over-optimizing chunk size, focus on:”
      ↩︎ Getting precision and context without one long window
      “Recent work from DeepMind (LIMIT) shows embeddings can fail to capture basic information in long contexts, making this a nuanced decision.”
      ↩︎ Exam trap 2
      “Larger chunks: More context preserved, but harder to pinpoint relevant information.”
      ↩︎ Checkpoint
    2. 2.
      “ensures that retrieved data fits in the LLM's context while minimizing the inclusion of distracting or irrelevant information”
      ↩︎ Why the maximum is a ceiling, not a target
      “Larger chunks may capture more context but can include irrelevant information or be computationally expensive.”
      ↩︎ Choosing chunk sizes to test
      “Larger embedding models generally perform better but require more computational resources.”
      ↩︎ Choosing chunk sizes to test
      “The optimal chunk size and method depends on the specific use case and the nature of the data being processed.”
      ↩︎ Choosing chunk sizes to test

    Also cited

    Spotted a mistake, or was something unclear? Tell us.