CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 22/56

    Embedding Model Context Length: Fitting Documents and Queries

    Select an embedding model context length based on source documents, expected queries, and optimization strategy

    9 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what an embedding model's token limit is and what happens to text that exceeds it
    • Compare the embedding windows of databricks-gte-large-en and databricks-bge-large-en and say which source documents each can embed without truncation
    • Account for expected queries when choosing an embedding model, including the rule that queries and chunks share one model

    Key concept

    Embedding window (max tokens) — Every embedding model has a fixed number of tokens it can turn into one vector. Text past that limit does not get embedded, so the chunks you send must fit inside the window, and that fact shapes both model choice and chunking.

    1.What an embedding model's context length is

    In a RAG pipeline, the embedding step sits between chunking and indexing. An embedding model takes each text chunk and turns it into a dense vector that captures what the chunk means. Retrieval then works by comparing the query's vector with those chunk vectors. The Databricks data-pipeline guide lists the factors to weigh when choosing that model: model choice, model size, fine-tuning, and max tokens. This lesson is about the last one: how many tokens the model can embed in a single call, also called its context length or embedding window.

    This is why context length is a selection criterion and not a footnote. If a chunk's key fact sits in its last few hundred tokens and the model cuts the chunk off before that point, the fact isn't in the vector. A query asking about it has nothing to match. The guide gives a concrete figure: "For example, bge-large-en-v1.5 has a maximum token limit of 512." Its advice is direct: know the maximum token limit for your chosen embedding model before you design anything that feeds it.

    The two embedding endpoints served by Databricks Foundation Model APIs show how far apart these limits can be. Both produce 1024-dimension vectors. Their windows differ by a factor of sixteen.

    Embedding windows of the two Foundation Model API embedding endpoints
    EndpointEmbedding windowVector sizeLongest chunk it embeds without truncation
    databricks-gte-large-en8192 tokens1024 dimensionsUp to 8192 tokens
    databricks-bge-large-en512 tokens1024 dimensionsUp to 512 tokens

    Checkpoint 1 of 4· Check yourself

    A team wants to embed chunks of about 2,000 tokens without losing any of their text. Which Databricks-served embedding endpoint can take those chunks whole?

    Sources12

    2.Sizing against your source documents

    The first input to the decision is the corpus itself. Source documents come in many forms. The guide's customer-support example lists knowledge base documents, FAQs, product manuals and specifications, and troubleshooting guides. After parsing, most of these are far longer than any embedding window, so they have to be split into chunks. Chunking exists partly to keep each piece inside the embedding model's limit. The Databricks AI Search example notebooks say so directly: "Chunking the sample dataset helps you avoid exceeding the context limit of the embedding model."

    The GTE example notebook shows the simplest form of this. It counts tokens with tiktoken and cuts each Wikipedia article into slices no larger than the model's window, so nothing gets truncated at embedding time.

    Token-based chunking capped at GTE's 8192-token context length (the comment's typo is in the source)python
    # The GTE model has been trained on a max context lenth of 8192 tokens.
    max_chunk_tokens = 8192
    encoding = tiktoken.get_encoding("cl100k_base")
    
    def chunk_text(text):
        # Encode and then decode within the UDF
        tokens = encoding.encode(text)
        chunks = []
        while tokens:
            chunk_tokens = tokens[:max_chunk_tokens]
            chunk_text = encoding.decode(chunk_tokens)
            chunks.append(chunk_text)
            tokens = tokens[max_chunk_tokens:]
        return chunks

    The window and the chunk size limit each other. If your documents divide naturally into long, self-contained units, such as full troubleshooting procedures or manual sections, and you want each unit embedded as one vector, you need a model whose window is at least that long. GTE's 8192 tokens allows that. BGE's 512 tokens does not. If your content already breaks into short paragraphs or FAQ entries, a 512-token window may be enough. Context length is still only one factor. The same guide warns that benchmarks may not reflect your data: "It's crucial to select a model that has been trained on similar data."

    Checkpoint 2 of 4· Exam question

    A team is building a RAG system over lengthy legal contracts averaging 6,000 tokens per semantically chunked passage. Retrieval quality tests show many chunks are silently truncated before embedding, causing key clauses near the end of long chunks to be dropped from the vector representation. The team is currently using the `databricks-bge-large-en` embedding model. Which change best addresses the root cause?

    Sources31

    3.Accounting for expected queries

    No. A query is matched by comparing its vector with chunk vectors, so both sides must come from the same model. The guide states: "The retrieval query will be transformed at query time using the same embedding model used to embed chunks in the data pipeline." So the embedding model you pick for your documents also has to suit your queries. That includes its context length, since a query and any text added to it also have to fit in the window.

    How do you know what the queries will look like? The guide recommends involving domain experts and stakeholders from the start, because they "can provide insights into the types of queries that users are likely to submit." The kind of query should guide how you chunk. The AI Search retrieval quality guide labels 256-token chunks as "Better for precise fact retrieval". Its trade-off list says smaller chunks give better localization of specific information but may lose context. Larger chunks keep more context but make the relevant passage harder to pinpoint. Queries that look for one specific value suit small chunks. Questions that need surrounding explanation suit larger ones, and therefore a model whose window can hold them.

    Some models treat the query side differently. For databricks-bge-large-en, the model list notes that you may be able to improve the performance of your retrieval system by including an instruction parameter. The BGE authors suggest a specific instruction for query embeddings. The same note adds that the performance impact depends on the domain, so test it rather than assume it helps.

    Checkpoint 3 of 4· Check yourself

    A RAG app embeds its document chunks with databricks-gte-large-en. Which model should embed incoming user queries at retrieval time?

    Checkpoint 4 of 4· Exam question

    A support-ticket search application embeds customer FAQ entries and short troubleshooting snippets that average 180 tokens each, with no snippet exceeding 400 tokens. The team wants to minimize serving cost and query latency while keeping retrieval accuracy on these short passages. Which embedding model choice best fits this optimization strategy?

    Sources142

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.If a chunk is longer than the embedding model's token limit, the endpoint returns an error, so oversized chunks are easy to catch.Why is that wrong?

      Oversized chunks are truncated, and the content past the limit is lost from the vector. Check chunk lengths against the model's max tokens before embedding.

      Covered in What an embedding model's context length is

    2. 2.Queries are short, so they can be embedded with a different, cheaper embedding model than the one used for the documents.Why is that wrong?

      At query time the query must be embedded with the same model that embedded the chunks. Otherwise the vectors cannot be meaningfully compared.

      Covered in Accounting for expected queries

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “For example, bge-large-en-v1.5 has a maximum token limit of 512.”
      ↩︎ What an embedding model's context length is
      “It's crucial to select a model that has been trained on similar data.”
      ↩︎ Sizing against your source documents
      “They can provide insights into the types of queries that users are likely to submit”
      ↩︎ Accounting for expected queries
      “If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
      ↩︎ Key concept
      “If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
      ↩︎ Exam trap 1
      “The retrieval query will be transformed at query time using the same embedding model used to embed chunks in the data pipeline.”
      ↩︎ Exam trap 2
      “If you pass chunks that exceed this limit, they will be truncated, potentially losing important information.”
      ↩︎ Prediction
      “The retrieval query will be transformed at query time using the same embedding model used to embed chunks in the data pipeline.”
      ↩︎ Checkpoint
    2. 2.
      “an embedding window of 512 tokens”
      ↩︎ What an embedding model's context length is
      “you may be able to improve the performance of your retrieval system by including an instruction parameter”
      ↩︎ Accounting for expected queries
      “map any text to a 1024-dimension embedding vector and an embedding window of 8192 tokens”
      ↩︎ Checkpoint
    3. 3.
      “Chunking the sample dataset helps you avoid exceeding the context limit of the embedding model.”
      ↩︎ Sizing against your source documents
    4. 4.
      “Smaller chunks: Better localization of specific information, but may lose context.”
      ↩︎ Accounting for expected queries

    Continue to page 2 of 2

    Chunk Size vs Embedding Context Length: Optimization Strategy

    Spotted a mistake, or was something unclear? Tell us.