What you will be able to do
- List the workspace, table and privilege prerequisites for creating a Vector Search (AI Search) index
- Create a Standard or Storage Optimized endpoint with the Python SDK and explain how that choice limits what you can do with the index
- Choose between a Delta Sync Index and a Direct Vector Access Index, and between Databricks-computed and precomputed embeddings
- Pick a sync mode and bring a new index to the ONLINE state
Key concept
Endpoint + index — Vector search on Databricks has two parts. The endpoint is the serving compute that hosts and answers queries. The index is a Unity Catalog object built from a source table, and you attach it to an endpoint when you create it.
1.What has to be in place first
The exam guide says "Vector Search", but the current documentation calls the product Databricks AI Search. It is the same product under a new name. The SDK package is databricks-ai-search, its client class is AISearchClient, and the REST paths still start with /api/2.0/vector-search/. You need two things before you can run a single similarity query: an endpoint to serve the index and the index itself. You can build both from the UI, the Python SDK or the REST API.
The full prerequisite list is short. You need a Unity Catalog enabled workspace, serverless compute turned on, and, for standard endpoints, a source table with a change data feed. The source can be a Delta Lake table, a streaming table, or a managed Iceberg v3+ table. The index is created as a Unity Catalog object, so you need CREATE TABLE on the target schema. Rights to create and manage endpoints are a separate matter, controlled by endpoint access control lists. In a notebook, install the SDK with %pip install databricks-ai-search, restart Python, and import AISearchClient from databricks.ai_search.client.
Checkpoint 1 of 6· Check yourself
A data engineer can read the source table but gets a permission error when creating an index in main.rag. Which privilege is missing?
The index is a Unity Catalog object created inside a schema, so creating it needs CREATE TABLE on that schema.
“To create an AI Search index, you must have CREATE TABLE privileges on the catalog schema where the index will be created.”Source: docs.databricks.com
Sources1
2.Creating the endpoint: Standard or Storage Optimized
In the UI, go to Compute → AI Search → Create endpoint, enter a name, and choose a Type. In code it is a single call. The type can only be set when the endpoint is created, and it affects almost everything you do later with indexes on that endpoint.
client.create_endpoint(
name="vector_search_endpoint_name",
endpoint_type="STANDARD" # or "STORAGE_OPTIMIZED"
)| Behaviour | STANDARD | STORAGE_OPTIMIZED |
|---|---|---|
| Sync modes for a Delta Sync Index | Continuous or Triggered | Triggered only |
| Source table change data feed | Required | Not listed as a requirement |
| Precomputed embedding dimension | No divisibility rule stated | Must be evenly divisible by 16 |
| Dedicated full-text (no-embedding) index | Not available | Available (Beta) |
| Query filter syntax | Python dictionary | SQL-like filter string |
| target_qps for high throughput | Supported | Not supported |
On a standard endpoint, target_qps provisions extra throughput. You pay for that capacity whether or not queries use it. Sizing and cost decisions belong to a separate objective. What matters here is that the type you pick now decides which sync mode, filter syntax and index subtype you can use later.
Checkpoint 2 of 6· Check yourself
A team sets up a standard endpoint with target_qps=500 for a demo that gets a few queries a day. What is the consequence?
target_qps reserves capacity in advance, and that reserved capacity is billed whatever the traffic. It is a standard-endpoint feature.
“Setting target_qps provisions additional capacity, which increases the cost of the endpoint.”Source: docs.databricks.com
3.Choosing the index type and the embedding source
The first decision is how the index is kept up to date. A Delta Sync Index follows a source table and updates incrementally as that table changes. A Direct Vector Access Index has no source table. You read and write vectors and metadata yourself through the REST API or SDK, and you cannot create one in the UI. A Delta Sync Index built without any embedding columns is a dedicated full-text index (Beta). That subtype is only available on storage-optimized endpoints with triggered sync.
For a Delta Sync Index you then decide where the embeddings come from. With Compute embeddings, you point at a text column and Databricks embeds it through a serving endpoint. For production on standard endpoints, the docs recommend databricks-qwen3-embedding-0-6b on provisioned throughput. With Use existing embeddings, you point at an array[float] column and supply its dimension. This choice is permanent: a self-managed embedding index cannot be converted to a Databricks-managed one later, so switching means building a new index.
Checkpoint 3 of 6· Match them up
Match each index option to how it behaves
Tap a term, then the definition that fits it.
The index type decides how data gets into the index. The embedding source decides who computes the vectors.
“This type of index cannot be created using the UI. You must use the REST API or the SDK.”Source: docs.databricks.com
There are a few practical details to watch. _id is a reserved column name, so rename any source column called _id first. The index name must use the three-level form <catalog>.<schema>.<name>. If you restrict Columns to index, keep in mind that only indexed columns can be returned in results or used as filters. If Databricks computes the embeddings, turn off *Scale to zero* on the embedding endpoint, because a cold endpoint can make the first query time out. CPU embedding endpoints are only suitable for small datasets and tests.
Checkpoint 4 of 6· Fill the gap
Which client method completes this sample that builds an index synced from a source table?
index = vsc. ? (
endpoint_name=ai_search_endpoint_name,
source_table_name=source_table_fullname,
index_name=vs_index_fullname,
pipeline_type='TRIGGERED',
primary_key="id",
embedding_source_column="text",
embedding_model_endpoint_name=embedding_model_endpoint
)The call has a source_table_name and an embedding_source_column, which means a Delta Sync Index with Databricks-computed embeddings. A Direct Vector Access Index has no source table.
Source: docs.databricks.comSources1
4.Sync mode, triggering a sync, and waiting for ONLINE
A Delta Sync Index uses one of two pipeline types. Continuous sync keeps the index within seconds of the source, but it costs more because a compute cluster runs a streaming pipeline the whole time. With Triggered sync, nothing runs until you start a sync through the SDK or REST API, for example index.sync(), or databricks vector-search-indexes sync-index from the CLI. On standard endpoints, both modes process only the rows that changed. On storage-optimized endpoints only Triggered is allowed, and each sync partially rebuilds the index.
No. Sync only applies to Delta Sync Indexes. A Direct Vector Access Index is updated by writing (upserting) vectors to it directly.
Building a new index takes a while. The SDK sample polls describe() until detailed_state starts with ONLINE, and only then sends queries:
import time
while not index.describe().get('status').get('detailed_state').startswith('ONLINE'):
print("Waiting for index to be ONLINE...")
time.sleep(5)
print("Index is ONLINE")Checkpoint 5 of 6· Put it in order
Put the SDK workflow for a Databricks-embedded Delta Sync Index in order
- 1.Run similarity_search() on the index
- 2.Poll describe() until detailed_state is ONLINE
- 3.Create the endpoint with create_endpoint()
- 4.Write the source Delta table with delta.enableChangeDataFeed set to true
- 5.Call create_delta_sync_index() and pass that endpoint_name
The index has to name an existing endpoint, and on a standard endpoint the source table needs a change data feed. Queries are only sent once the index is online.
“An AI Search endpoint. This endpoint serves the AI Search index.”Source: docs.databricks.com
Checkpoint 6 of 6· Exam question
A team ingests product documentation into a Unity Catalog Delta table that receives new rows several times a day. They want a vector search index that stays automatically consistent with the table without writing custom code to push updates. Which indexing approach should they use?
Correct answer: A — Create a Delta Sync index on the source Delta table so the AI Search index incrementally syncs new and changed rows automatically.
- A. A Delta Sync index is built directly on a source Delta table and Databricks incrementally syncs inserts, updates, and deletes automatically, which is exactly what the team needs without custom code.
- B. A Direct Vector Access index does support upsert calls, but that requires the team to write and schedule application code to push each change, which contradicts their goal of avoiding custom update logic.
- C. A Direct Vector Access index does not automatically read Change Data Feed; vectors and metadata must be pushed explicitly through the upsert or delete API, so there is no automatic population from table changes.
- D. A triggered pipeline only syncs when a `sync()` call is explicitly made on the index; it does not pick up changes on its own, so this configuration would leave the index stale.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A storage-optimized endpoint can run a Continuous sync so the index stays fresh within seconds.Why is that wrong?
Storage-optimized endpoints only allow Triggered sync. Continuous sync requires a standard endpoint.
Covered in Sync mode, triggering a sync, and waiting for ONLINE
2.Any index type can be created from Catalog Explorer with Create → Vector search index.Why is that wrong?
A Direct Vector Access Index can only be created through the REST API or the SDK.
3.You can start with precomputed embeddings and later switch the same index to Databricks-computed embeddings.Why is that wrong?
The conversion is not possible. You would need to create a new index and recompute the embeddings.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Databricks AI Search was formerly known as Databricks Vector Search.”
↩︎ What has to be in place first“For standard endpoints, the source table must use a change data feed.”
↩︎ What has to be in place first“For storage-optimized endpoints, the embedding dimension must be evenly divisible by 16.”
↩︎ Creating the endpoint: Standard or Storage Optimized“Delta Sync Index automatically syncs with a source table, incrementally updating the index as the underlying data in the source table changes.”
↩︎ Choosing the index type and the embedding source“The column name _id is reserved.”
↩︎ Choosing the index type and the embedding source“Databricks recommends that you remove the default selection of Scale to zero.”
↩︎ Choosing the index type and the embedding source“Continuous keeps the index in sync with seconds of latency. However, it has a higher cost associated with it”
↩︎ Sync mode, triggering a sync, and waiting for ONLINE“With Triggered sync mode, you use the Python SDK or the REST API to start the sync.”
↩︎ Sync mode, triggering a sync, and waiting for ONLINE“For storage-optimized endpoints, only Triggered sync mode is supported.”
↩︎ Exam trap 1“This type of index cannot be created using the UI. You must use the REST API or the SDK.”
↩︎ Exam trap 2“To create an AI Search index, you must have CREATE TABLE privileges on the catalog schema where the index will be created.”
↩︎ Checkpoint“Setting target_qps provisions additional capacity, which increases the cost of the endpoint.”
↩︎ Checkpoint“This type of index cannot be created using the UI. You must use the REST API or the SDK.”
↩︎ Checkpoint - 2.
“You specify the endpoint type when you create the endpoint.”
↩︎ Creating the endpoint: Standard or Storage Optimized“An AI Search endpoint. This endpoint serves the AI Search index.”
↩︎ Key concept“It is not possible to convert a self-managed embedding index to a Databricks-managed index.”
↩︎ Exam trap 3“An AI Search endpoint. This endpoint serves the AI Search index.”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/dev-tools/cli/reference/vector-search-indexes-commandsOfficial docs
“Name of the vector index to synchronize. Must be a Delta Sync Index.”
↩︎ Sync mode, triggering a sync, and waiting for ONLINE