CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 4 · Lesson 37/56

    Vector Search Sync Modes, High QPS and Query Latency Tuning

    Configure vector search for a particular solution based on number of embeddings, update frequency, latency, and cost requirements.

    9 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Pick a sync mode and index type that match how often the source data changes and how much freshness is worth
    • Decide when to set target_qps on an endpoint and what it costs
    • Configure the query path (embeddings, query type, num_results, authentication) for low latency

    1.Match the sync mode to update frequency

    How often your source data changes, and how quickly the index has to show those changes, decides how the index is kept up to date. There are two index types. A Delta Sync Index syncs from a source table (a Delta table, a streaming table, or a managed Iceberg v3+ table) and updates incrementally as that table changes. A Direct Vector Access Index puts you in charge: you write vectors and metadata yourself through the REST API or Python SDK, and it cannot be created in the UI.

    For Delta Sync indexes you then choose a sync mode.

    Update options compared
    OptionHow the index updatesFreshness and cost
    Delta Sync, TriggeredYou start each sync with the Python SDK or REST APIChanges appear when you sync; the most cost-effective mode; the only mode on storage-optimized endpoints
    Delta Sync, ContinuousA streaming pipeline keeps the index in sync automaticallySeconds of latency; higher cost because a compute cluster is provisioned
    Direct Vector AccessYour code writes vectors and metadata via the REST API or Python SDKFreshness depends on how often your code writes

    The endpoint type changes how syncing works. On standard endpoints, both Continuous and Triggered sync are incremental: only data changed since the last sync is processed. Standard endpoints also need a change data feed on the source table (row tracking turns one on automatically). On storage-optimized endpoints, only Triggered sync is supported, and every sync partially rebuilds the index. For managed embeddings, rows that haven't changed reuse their existing embeddings rather than being recomputed. So a source that changes constantly and must show up within seconds rules out storage-optimized. For large, slow-changing corpora, Triggered sync on storage-optimized fits well.

    Checkpoint 1 of 4· Fill the gap

    This index is created on a storage-optimized endpoint. Which value must pipeline_type have?

    index = client.create_delta_sync_index(
      endpoint_name="vector_search_demo_endpoint",
      source_table_name="vector_search_demo.vector_search.en_wiki",
      index_name="vector_search_demo.vector_search.en_wiki_index",
      pipeline_type=" ? ",
      primary_key="id",
      embedding_dimension=1024,
      embedding_vector_column="text_vector"
    )

    Sources123

    2.Provision throughput with target_qps

    By default, a standard endpoint handles 20–200 QPS depending on index size. Real-time applications such as search bars, recommendation systems and entity matching often need 100–1000+. On standard endpoints only, you can set target_qps, and Databricks provisions infrastructure to match that throughput as closely as it can. It is best-effort, not guaranteed. Consider it when you need more than 50 QPS sustained, when you get 429 (Too Many Requests) errors under normal load, or when latency degrades as traffic ramps up.

    Creating a standard endpoint with a target QPSpython
    client.create_endpoint(
        name="vector_search_endpoint_name",
        endpoint_type="STANDARD",
        target_qps=500,  # target QPS for high-throughput workloads
    )

    This raises cost directly. The extra capacity is billed whether or not traffic arrives, and nothing autoscales: you set the number from expected traffic, and traffic above it returns 429s. The setting also covers only the optimized query route, which means service principal OAuth plus the index URL. PAT traffic and traffic to the workspace query URL are capped at a few tens of QPS. With managed embeddings, the embedding model serving endpoint can become a second bottleneck for text queries. In that case, add model serving capacity, use provisioned throughput, or switch to self-managed embeddings.

    Checkpoint 2 of 4· Put it in order

    Put these steps for bringing a high-QPS endpoint into production in order

    1. 1.Wait until scaling_info.state changes from SCALING_CHANGE_IN_PROGRESS to SCALING_CHANGE_APPLIED
    2. 2.Send queries to the index URL using a service principal OAuth token
    3. 3.Set target_qps on the standard endpoint with create_endpoint() or update_endpoint()

    Checkpoint 3 of 4· Exam question

    A data science team already generates embeddings for their documents using a custom fine-tuned model in an external pipeline, and they store the resulting vectors alongside the source rows in a Delta table. They want the vector search index to sync automatically from that Delta table without Databricks recomputing the vectors. Which index configuration fits this requirement?

    Sources4

    3.Tune the query path for latency

    Endpoint type and capacity set the ceiling. How each query is made decides how close you get to it.

    Embeddings. With managed embeddings, Databricks embeds the query text by calling a model serving endpoint, which adds latency, and more again if the model is external. Self-managed embeddings send a precomputed query_vector, so no model is called at query time and retrieval is fastest. If you use managed embeddings, do not let the embedding endpoint scale to zero: a cold start can delay responses by minutes or make them fail. CPU endpoints suit only small datasets and tests; larger ones need GPU. You can also give queries their own endpoint for the same model, pairing a high-throughput endpoint for ingestion with a low-latency one for queries:

    A separate low-latency query endpoint via model_endpoint_name_for_querypython
    index = client.create_delta_sync_index(
      endpoint_name="vector_search_demo_endpoint",
      source_table_name="vector_search_demo.vector_search.en_wiki",
      index_name="vector_search_demo.vector_search.en_wiki_index",
      pipeline_type="TRIGGERED",
      primary_key="id",
      embedding_source_column="text",
      embedding_model_endpoint_name="databricks-qwen3-embedding-0-6b", # Used for ingestion, and also for querying unless model_endpoint_name_for_query is specified.
      model_endpoint_name_for_query="qwen3-embedding-low-latency"      # Optional. A separate endpoint serving the same model, used only for querying.
    )

    Query type and result count. Use ANN wherever you can, since it is the most compute-efficient and supports the highest QPS. Hybrid search improves recall when exact keywords such as SKUs matter, but it uses about twice the resources. Keep num_results between 10 and 100: increasing it 10x can double latency and cut QPS capacity by 3x.

    Client. In production, authenticate with a service principal OAuth token. PATs add hundreds of milliseconds and lower the QPS the endpoint can sustain. Create the index object once and reuse it for every query, instead of calling get_index() on each request. Use the Python SDK, which has built-in backoff and retry for 429s. If you call the REST API, implement exponential backoff with jitter yourself.

    Checkpoint 4 of 4· Check yourself

    A chatbot's retrieval step is too slow. It uses managed embeddings, hybrid search, num_results=500 and a PAT. Which change does the performance guide NOT support?

    Sources51

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A storage-optimized endpoint can use Continuous sync to keep a large index fresh within seconds.Why is that wrong?

      Storage-optimized endpoints support only Triggered sync, and every sync partially rebuilds the index. If you need freshness within seconds, use Continuous sync on a standard endpoint.

      Covered in Match the sync mode to update frequency

    2. 2.target_qps only costs money when traffic actually uses the extra capacity, and it works on any endpoint type.Why is that wrong?

      The provisioned capacity is billed whatever the traffic, and target_qps is available on standard endpoints only.

      Covered in Provision throughput with target_qps

    3. 3.Managed and self-managed embeddings have the same query latency once the index is built.Why is that wrong?

      Managed embeddings call a model serving endpoint at query time, which adds latency. Self-managed query vectors skip that call.

      Covered in Tune the query path for latency

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Continuous keeps the index in sync with seconds of latency. However, it has a higher cost associated with it”
      ↩︎ Match the sync mode to update frequency
      “For storage-optimized endpoints, every sync partially rebuilds the index.”
      ↩︎ Match the sync mode to update frequency
      “The user is responsible for updating this table using the REST API or the Python SDK.”
      ↩︎ Match the sync mode to update frequency
      “That lets you pair a high-throughput endpoint for ingestion with a low-latency endpoint for queries.”
      ↩︎ Tune the query path for latency
      “For storage-optimized endpoints, only Triggered sync mode is supported.”
      ↩︎ Exam trap 1
    2. 2.
      “For standard endpoints, the source table must use a change data feed.”
      ↩︎ Match the sync mode to update frequency
    3. 3.
      “Triggered Sync: You call the API or Python SDK to trigger an index update. This is the most cost-effective option.”
      ↩︎ Match the sync mode to update frequency
      “If near real-time updates with seconds of latency are not critical, consider using Triggered Sync to reduce costs.”
      ↩︎ Prediction
    4. 4.
      “On standard endpoints only, you can set a target QPS.”
      ↩︎ Provision throughput with target_qps
      “No autoscaling: You must set target QPS manually based on expected traffic. If traffic exceeds the provisioned level, 429 errors occur.”
      ↩︎ Provision throughput with target_qps
      “Your application requires more than 50 QPS of sustained throughput.”
      ↩︎ Provision throughput with target_qps
      “You are charged for this additional capacity regardless of actual query traffic.”
      ↩︎ Exam trap 2
      “After the endpoint scaling state is SCALING_CHANGE_APPLIED, send queries to the index URL using a service principal OAuth token.”
      ↩︎ Checkpoint
    5. 5.
      “Use ANN queries whenever possible. They are the most compute-efficient and support the highest QPS.”
      ↩︎ Tune the query path for latency
      “keep num_results in the range of 10–100 unless your application specifically requires more.”
      ↩︎ Tune the query path for latency
      “Avoid using personal access tokens (PATs), as they introduce network overhead, add hundreds of milliseconds of latency”
      ↩︎ Tune the query path for latency
      “At query time, the query text is passed to a model serving endpoint to generate the embedding, which adds latency.”
      ↩︎ Exam trap 3
      “For real-time production use cases, avoid model endpoints that scale to zero.”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.