What you will be able to do
- Pick a sync mode and index type that match how often the source data changes and how much freshness is worth
- Decide when to set target_qps on an endpoint and what it costs
- Configure the query path (embeddings, query type, num_results, authentication) for low latency
1.Match the sync mode to update frequency
How often your source data changes, and how quickly the index has to show those changes, decides how the index is kept up to date. There are two index types. A Delta Sync Index syncs from a source table (a Delta table, a streaming table, or a managed Iceberg v3+ table) and updates incrementally as that table changes. A Direct Vector Access Index puts you in charge: you write vectors and metadata yourself through the REST API or Python SDK, and it cannot be created in the UI.
For Delta Sync indexes you then choose a sync mode.
| Option | How the index updates | Freshness and cost |
|---|---|---|
| Delta Sync, Triggered | You start each sync with the Python SDK or REST API | Changes appear when you sync; the most cost-effective mode; the only mode on storage-optimized endpoints |
| Delta Sync, Continuous | A streaming pipeline keeps the index in sync automatically | Seconds of latency; higher cost because a compute cluster is provisioned |
| Direct Vector Access | Your code writes vectors and metadata via the REST API or Python SDK | Freshness depends on how often your code writes |
The endpoint type changes how syncing works. On standard endpoints, both Continuous and Triggered sync are incremental: only data changed since the last sync is processed. Standard endpoints also need a change data feed on the source table (row tracking turns one on automatically). On storage-optimized endpoints, only Triggered sync is supported, and every sync partially rebuilds the index. For managed embeddings, rows that haven't changed reuse their existing embeddings rather than being recomputed. So a source that changes constantly and must show up within seconds rules out storage-optimized. For large, slow-changing corpora, Triggered sync on storage-optimized fits well.
Checkpoint 1 of 4· Fill the gap
This index is created on a storage-optimized endpoint. Which value must pipeline_type have?
index = client.create_delta_sync_index(
endpoint_name="vector_search_demo_endpoint",
source_table_name="vector_search_demo.vector_search.en_wiki",
index_name="vector_search_demo.vector_search.en_wiki_index",
pipeline_type=" ? ",
primary_key="id",
embedding_dimension=1024,
embedding_vector_column="text_vector"
)Storage-optimized endpoints support only Triggered sync. STANDARD and STORAGE_OPTIMIZED are endpoint types, not pipeline types.
Source: docs.databricks.com2.Provision throughput with target_qps
By default, a standard endpoint handles 20–200 QPS depending on index size. Real-time applications such as search bars, recommendation systems and entity matching often need 100–1000+. On standard endpoints only, you can set target_qps, and Databricks provisions infrastructure to match that throughput as closely as it can. It is best-effort, not guaranteed. Consider it when you need more than 50 QPS sustained, when you get 429 (Too Many Requests) errors under normal load, or when latency degrades as traffic ramps up.
client.create_endpoint(
name="vector_search_endpoint_name",
endpoint_type="STANDARD",
target_qps=500, # target QPS for high-throughput workloads
)This raises cost directly. The extra capacity is billed whether or not traffic arrives, and nothing autoscales: you set the number from expected traffic, and traffic above it returns 429s. The setting also covers only the optimized query route, which means service principal OAuth plus the index URL. PAT traffic and traffic to the workspace query URL are capped at a few tens of QPS. With managed embeddings, the embedding model serving endpoint can become a second bottleneck for text queries. In that case, add model serving capacity, use provisioned throughput, or switch to self-managed embeddings.
Checkpoint 2 of 4· Put it in order
Put these steps for bringing a high-QPS endpoint into production in order
- 1.Wait until scaling_info.state changes from SCALING_CHANGE_IN_PROGRESS to SCALING_CHANGE_APPLIED
- 2.Send queries to the index URL using a service principal OAuth token
- 3.Set target_qps on the standard endpoint with create_endpoint() or update_endpoint()
Capacity is provisioned first and is usable once the state is SCALING_CHANGE_APPLIED. Queries must then use the optimized route to get it.
“After the endpoint scaling state is SCALING_CHANGE_APPLIED, send queries to the index URL using a service principal OAuth token.”Source: docs.databricks.com
Checkpoint 3 of 4· Exam question
A data science team already generates embeddings for their documents using a custom fine-tuned model in an external pipeline, and they store the resulting vectors alongside the source rows in a Delta table. They want the vector search index to sync automatically from that Delta table without Databricks recomputing the vectors. Which index configuration fits this requirement?
Correct answer: A — A Self-Managed embeddings index that reads the precomputed vectors from the designated `array[float]` column during Delta sync.
- A. A Self-Managed embeddings index still uses Delta Sync against the source table but reads precomputed vectors from an `array[float]` column instead of recomputing them, which matches a team that already produces embeddings externally and wants automatic syncing.
- B. A Databricks-Managed embeddings index computes vectors itself using a specified embedding model, so it would recompute vectors from raw text rather than reuse the team's existing custom-model embeddings.
- C. A Direct Vector Access index does not sync from a Delta table's change feed at all; it requires the application to explicitly read and write vectors and metadata itself, so it does not provide automatic syncing.
- D. Disabling sync on a Databricks-Managed index would freeze the index at its initial state and prevent it from picking up new rows added to the Delta table, defeating the goal of ongoing automatic updates.
Sources4
3.Tune the query path for latency
Endpoint type and capacity set the ceiling. How each query is made decides how close you get to it.
Embeddings. With managed embeddings, Databricks embeds the query text by calling a model serving endpoint, which adds latency, and more again if the model is external. Self-managed embeddings send a precomputed query_vector, so no model is called at query time and retrieval is fastest. If you use managed embeddings, do not let the embedding endpoint scale to zero: a cold start can delay responses by minutes or make them fail. CPU endpoints suit only small datasets and tests; larger ones need GPU. You can also give queries their own endpoint for the same model, pairing a high-throughput endpoint for ingestion with a low-latency one for queries:
index = client.create_delta_sync_index(
endpoint_name="vector_search_demo_endpoint",
source_table_name="vector_search_demo.vector_search.en_wiki",
index_name="vector_search_demo.vector_search.en_wiki_index",
pipeline_type="TRIGGERED",
primary_key="id",
embedding_source_column="text",
embedding_model_endpoint_name="databricks-qwen3-embedding-0-6b", # Used for ingestion, and also for querying unless model_endpoint_name_for_query is specified.
model_endpoint_name_for_query="qwen3-embedding-low-latency" # Optional. A separate endpoint serving the same model, used only for querying.
)Query type and result count. Use ANN wherever you can, since it is the most compute-efficient and supports the highest QPS. Hybrid search improves recall when exact keywords such as SKUs matter, but it uses about twice the resources. Keep num_results between 10 and 100: increasing it 10x can double latency and cut QPS capacity by 3x.
Client. In production, authenticate with a service principal OAuth token. PATs add hundreds of milliseconds and lower the QPS the endpoint can sustain. Create the index object once and reuse it for every query, instead of calling get_index() on each request. Use the Python SDK, which has built-in backoff and retry for 429s. If you call the REST API, implement exponential backoff with jitter yourself.
Checkpoint 4 of 4· Check yourself
A chatbot's retrieval step is too slow. It uses managed embeddings, hybrid search, num_results=500 and a PAT. Which change does the performance guide NOT support?
Scale-to-zero causes cold starts that can delay queries by minutes or make them fail. The other three are all recommended latency fixes.
“For real-time production use cases, avoid model endpoints that scale to zero.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A storage-optimized endpoint can use Continuous sync to keep a large index fresh within seconds.Why is that wrong?
Storage-optimized endpoints support only Triggered sync, and every sync partially rebuilds the index. If you need freshness within seconds, use Continuous sync on a standard endpoint.
Covered in Match the sync mode to update frequency
2.target_qps only costs money when traffic actually uses the extra capacity, and it works on any endpoint type.Why is that wrong?
The provisioned capacity is billed whatever the traffic, and target_qps is available on standard endpoints only.
Covered in Provision throughput with target_qps
3.Managed and self-managed embeddings have the same query latency once the index is built.Why is that wrong?
Managed embeddings call a model serving endpoint at query time, which adds latency. Self-managed query vectors skip that call.
Covered in Tune the query path for latency
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Continuous keeps the index in sync with seconds of latency. However, it has a higher cost associated with it”
↩︎ Match the sync mode to update frequency“For storage-optimized endpoints, every sync partially rebuilds the index.”
↩︎ Match the sync mode to update frequency“The user is responsible for updating this table using the REST API or the Python SDK.”
↩︎ Match the sync mode to update frequency“That lets you pair a high-throughput endpoint for ingestion with a low-latency endpoint for queries.”
↩︎ Tune the query path for latency“For storage-optimized endpoints, only Triggered sync mode is supported.”
↩︎ Exam trap 1 - 2.
“For standard endpoints, the source table must use a change data feed.”
↩︎ Match the sync mode to update frequency - 3.
“Triggered Sync: You call the API or Python SDK to trigger an index update. This is the most cost-effective option.”
↩︎ Match the sync mode to update frequency“If near real-time updates with seconds of latency are not critical, consider using Triggered Sync to reduce costs.”
↩︎ Prediction - 4.
“On standard endpoints only, you can set a target QPS.”
↩︎ Provision throughput with target_qps“No autoscaling: You must set target QPS manually based on expected traffic. If traffic exceeds the provisioned level, 429 errors occur.”
↩︎ Provision throughput with target_qps“Your application requires more than 50 QPS of sustained throughput.”
↩︎ Provision throughput with target_qps“You are charged for this additional capacity regardless of actual query traffic.”
↩︎ Exam trap 2“After the endpoint scaling state is SCALING_CHANGE_APPLIED, send queries to the index URL using a service principal OAuth token.”
↩︎ Checkpoint - 5.
“Use ANN queries whenever possible. They are the most compute-efficient and support the highest QPS.”
↩︎ Tune the query path for latency“keep num_results in the range of 10–100 unless your application specifically requires more.”
↩︎ Tune the query path for latency“Avoid using personal access tokens (PATs), as they introduce network overhead, add hundreds of milliseconds of latency”
↩︎ Tune the query path for latency“At query time, the query text is passed to a model serving endpoint to generate the embedding, which adds latency.”
↩︎ Exam trap 3“For real-time production use cases, avoid model endpoints that scale to zero.”
↩︎ Checkpoint