What you will be able to do
- Name the two components you must create to use Vector Search (an endpoint and an index), the optional embedding model endpoint, and how they relate.
- Tell a Delta Sync Index from a Direct Vector Access Index, and pick between Databricks-managed embeddings, self-managed embeddings and a full-text-only index.
- Explain how similarity is computed (HNSW, L2 distance, normalization for cosine) and how hybrid search combines ANN and BM25 with Reciprocal Rank Fusion.
- Choose between the ANN, full-text, hybrid and reranked query types, and know the default query type and the filter syntax for each endpoint type.
Key concept
Endpoint serves, index holds — Vector Search is built from two objects you create. An index holds the embedded rows of a source table and is governed in Unity Catalog. An endpoint is the serving compute that hosts and answers queries against one or more indexes.
1.What Vector Search is and how it measures similarity
The exam guide calls this product Mosaic AI Vector Search. The current Databricks documentation calls it Databricks AI Search and notes that it was formerly Databricks Vector Search. Treat these as names for the same product. The SDK and docs use AISearchClient, but the REST paths still read /api/2.0/vector-search/..., so expect both names on the exam.
Vector Search is the retrieval layer for AI applications such as RAG systems, recommender systems and image recognition. It works on embeddings, which are numeric representations of the meaning of text or images, produced by a model. A query is turned into an embedding (or arrives as one), and the service returns the stored entries whose embeddings are closest to it, along with the documents attached to them. Because it is built into the platform, it shares the platform's governance: indexes live in Unity Catalog, and endpoints are controlled with access control lists.
Search is approximate, not exhaustive. Vector Search uses the Hierarchical Navigable Small World (HNSW) algorithm for approximate nearest neighbor (ANN) search, and L2 distance as its similarity metric. The cosine point from the predict above is the detail most likely to be tested. Normalization does not happen automatically: if your design depends on cosine similarity, you normalize the vectors before they reach the index.
Pure vector search has a known weakness. Embeddings capture meaning well but can blur exact tokens such as part numbers or error codes. So Vector Search also supports hybrid keyword-similarity search. A keyword search scores documents with Okapi BM25 across text and string columns, an ANN search runs alongside it, and the two ranked lists are merged with Reciprocal Rank Fusion (RRF). RRF rescores each document by its rank in each list (with rrf_param set to 60), normalizes the result so the best possible score is 1, and returns the documents with the highest combined scores.
Checkpoint 1 of 9· Check yourself
A RAG assistant for a parts catalogue often misses documents when users type exact SKUs, but handles conceptual questions well. Which Vector Search capability targets this gap?
Hybrid search adds exact keyword matching to semantic similarity. The docs call this out for keywords such as SKUs, which pure similarity search handles poorly.
“useful in RAG applications where source data has unique keywords such as SKUs or identifiers that are not well suited to pure similarity search”Source: docs.databricks.com
Sources1
2.The components: endpoint, index and embedding model endpoint
Using Vector Search means creating two things, plus an optional third.
1. A Vector Search endpoint. This is the compute that serves indexes. You can manage it through the UI (Compute → AI Search tab), the Python SDK or the REST API. It scales up automatically with index size and with the number of concurrent requests, and scales down when an index is deleted.
2. A Vector Search index. This is the searchable structure built from a source table: a Delta Lake table, a streaming table or a managed Iceberg v3+ table. It stores the embedded data together with metadata, supports real-time ANN queries, and is a Unity Catalog object. Its name therefore uses the three-level <catalog>.<schema>.<name> form.
3. An embedding model serving endpoint (optional). You only need this if Databricks computes the embeddings for you. It can be a pre-configured Foundation Model APIs endpoint or a model serving endpoint that you create for an embedding model of your choice.
You choose the endpoint type when you create it, and that choice limits what its indexes can do later:
| Endpoint type | Capacity | Indexing and latency | Constraints |
|---|---|---|---|
| STANDARD | 320 million vectors at dimension 768 | Standard query latency; high QPS available for sustained throughput | Source table must use a change data feed |
| STORAGE_OPTIMIZED | Over one billion vectors at dimension 768 | 10-20x faster indexing; queries about 250 msec slower | Triggered sync only; filters use a SQL-like string |
# The following line automatically generates a PAT Token for authentication
client = AISearchClient()
# The following line uses the service principal token for authentication
# client = AISearchClient(service_principal_client_id=<CLIENT_ID>,service_principal_client_secret=<CLIENT_SECRET>)
client.create_endpoint(
name="vector_search_endpoint_name",
endpoint_type="STANDARD" # or "STORAGE_OPTIMIZED"
)The surrounding requirements are part of the component picture:
- Workspace: Unity Catalog must be enabled, and so must serverless compute.
- Source table: for standard endpoints it needs a change data feed. Tables with row tracking enabled get one automatically.
- Permissions: creating an index requires CREATE TABLE on the target schema. Creating and managing endpoints is controlled through endpoint ACLs.
- Authentication: two modes are supported, service principals and personal access tokens (PATs). In a notebook, the SDK generates a PAT for you. For production, Databricks recommends service principals, which can also be up to 100 msec faster per query.
- Security: data is encrypted at rest (AES-256) and in transit (TLS 1.2+).
Checkpoint 2 of 9· Check yourself
An engineer creates a Delta Sync Index on a standard endpoint. The source Delta table has neither a change data feed nor row tracking. What must change?
On standard endpoints a change data feed on the source table is a prerequisite. Row tracking satisfies it automatically.
“For standard endpoints, the source table must use a change data feed.”Source: docs.databricks.com
Checkpoint 3 of 9· Exam question
A data engineering team is building a RAG index over an internal knowledge base that is refreshed only twice a day by a batch ETL job. Cost efficiency matters more than index freshness, and the team does not need a continuously running sync process. Which AI Search endpoint and sync configuration best fits these requirements?
Correct answer: C — Create the index on a storage-optimized endpoint with triggered sync, invoking a sync operation after each ETL batch completes
- A. Continuous sync keeps provisioned compute running to maintain near real-time latency, which costs more than needed when the source only changes twice a day. It is a mismatch between the required freshness and the ongoing compute expense of continuous sync.
- B. Change Data Feed is a requirement for keeping a Delta Sync index synchronized with its source table on a standard endpoint, so disabling it would break the sync mechanism rather than just save storage. This configuration would leave the index unable to detect changes at all.
- C. A storage-optimized endpoint only supports triggered sync, which matches a workload that updates a couple of times a day, and triggering the sync right after the ETL batch keeps the index current without paying for always-on compute. This combination minimizes cost while still refreshing the index on the required cadence.
- D. Direct Vector Access indexes require manually computing and upserting vectors via the REST API or SDK, which adds unnecessary custom engineering when the built-in Delta Sync mechanism already handles this pattern with a managed index type. It also loses the automatic synchronization that a Delta Sync index provides.
Sources1
3.Index types, embedding sources and sync modes
Before creating an index you make two independent decisions: how the index stays current (the index type) and where the embeddings come from.
Index type.
- A Delta Sync Index stays tied to a source table and updates incrementally as that table changes.
- A Direct Vector Access Index has no source table. You write vectors and metadata into it yourself with upsert and delete calls through the SDK or REST API. This type cannot be created in the UI.
Embedding source.
- Databricks-managed embeddings: the source holds text. Databricks computes embeddings with the model you name, and computes the query embedding at search time too. Optionally, it saves the computed embeddings to a Unity Catalog table named after the index with _writeback_table appended. That table is read-only and is deleted along with the index.
- Self-managed embeddings: the source already holds an array<float> column of vectors, and you give its dimension. Queries then pass query_vector instead of query_text.
- No embeddings (Beta): a dedicated full-text index for BM25 keyword search only, available only on storage-optimized endpoints.
The embedding source is permanent. You cannot convert a self-managed index into a managed one; you would have to create a new index and recompute the embeddings.
index = client.create_delta_sync_index(
endpoint_name="vector_search_demo_endpoint",
source_table_name="vector_search_demo.vector_search.en_wiki",
index_name="vector_search_demo.vector_search.en_wiki_index",
pipeline_type="TRIGGERED",
primary_key="id",
embedding_dimension=1024,
embedding_vector_column="text_vector"
)pipeline_type sets the sync mode of a Delta Sync Index:
- Continuous keeps the index within seconds of the source table. It costs more because a compute cluster runs a streaming sync pipeline.
- Triggered syncs only when you call index.sync(), click Sync now in Catalog Explorer, or call the REST sync endpoint.
On standard endpoints both modes update incrementally. On storage-optimized endpoints, Triggered is the only mode, each sync partially rebuilds the index, and embeddings for unchanged rows are reused.
A few structural rules come with the index:
- Every index has a primary key.
- columns_to_sync limits which source columns are indexed. The primary key and embedding column are always included, and only indexed columns can be returned or filtered on.
- The column name _id is reserved.
- The schema is fixed at creation. Adding or modifying source columns means building a new index.
Checkpoint 4 of 9· Match them up
Match each Vector Search component to its behaviour
Tap a term, then the definition that fits it.
A Delta Sync Index follows its source table, while a Direct Vector Access Index is written to directly by your own code. The full-text index skips embeddings entirely, and the writeback table stores managed embeddings.
“Direct Vector Access Index supports direct read and write of vectors and metadata.”Source: docs.databricks.com
The index schema is fixed when the index is created, so schema changes on the source table are not supported without a rebuild. The documented zero-downtime path is to build a second index on the new schema and switch traffic to it.
Checkpoint 5 of 9· Put it in order
Put the zero-downtime index rebuild steps in order
- 1.Delete the original index
- 2.Perform the schema change on your source table
- 3.After the new index is ready, switch traffic to the new index
- 4.Create a new index using the updated schema
The original index keeps serving until the new one is ready and traffic has moved. Only then is the original deleted, so there is no downtime.
“After the new index is ready, switch traffic to the new index.”Source: docs.databricks.com
Checkpoint 6 of 9· Exam question
A GenAI engineering team has already computed embeddings for their documents using a proprietary fine-tuned embedding model and stores the resulting vectors as a column in a Delta table alongside the document text. They want the index to automatically pick up new rows as the Delta table changes, without Databricks recomputing embeddings. Which index configuration should they use?
Correct answer: B — Create a Delta Sync index with self-managed embeddings, pointing the embedding vector column to the precomputed vectors already stored in the table
- A. Managed embeddings tell Databricks to generate vectors itself using a chosen embedding model, which would recompute embeddings and discard the team's proprietary fine-tuned vectors. This ignores the requirement to reuse the already-computed embeddings.
- B. Self-managed embeddings let a Delta Sync index reference an existing embedding column in the source table, so the index stays automatically synchronized as rows change while reusing the precomputed vectors exactly as required. This is the configuration built for exactly this scenario.
- C. A Direct Vector Access index bypasses automatic table synchronization entirely, requiring the team to manually push every vector and metadata update through the API, which contradicts the requirement for the index to pick up changes automatically. It adds operational overhead that a Delta Sync index avoids.
- D. Querying a vector column directly with SQL bypasses AI Search entirely and forgoes the approximate nearest neighbor indexing, scaling, and endpoint-based serving that a vector search index provides. Change Data Feed alone does not create a searchable index.
4.Querying: ANN, full-text, hybrid, reranking and filters
There are three ways to query an index: the Python SDK (similarity_search), the REST API, and the SQL vector_search() function. A user who does not own the index needs USE CATALOG, USE SCHEMA and SELECT on it. The query_type parameter selects the retrieval algorithm.
| Strategy | How it works | Best for |
|---|---|---|
| ANN (default) | Searches using vector embeddings to find semantically similar documents | Conceptual queries where meaning matters more than exact wording |
| Full-text (FULL_TEXT, Beta) | Keyword search that matches on exact terms | Proper nouns, product IDs, error codes, technical jargon |
| Hybrid | Combines ANN and full-text results using Reciprocal Rank Fusion (RRF) | General-purpose retrieval; the recommended starting point |
| Hybrid + reranker | Runs hybrid search, then re-scores results with a cross-encoder reranker model | Higher precision when latency allows |
Each strategy has limits worth knowing:
- Hybrid searches all text metadata columns by default and returns at most 200 results. - Full-text can return up to 10,000 keyword matches without using embeddings, on both endpoint types. - Reranker: an optional second pass that can sit on top of any strategy. A cross-encoder re-scores each retrieved result against the query. It typically improves quality by about 10% at the cost of extra latency. That suits RAG chatbots, where Databricks recommends trying it, more than high-throughput, low-latency search.
Which arguments a query passes depends on how the index was built. A managed-embedding index takes query_text, and a self-managed one takes query_vector.
Checkpoint 7 of 9· Fill the gap
Which query_type value runs ANN and keyword search in parallel and fuses them with RRF?
results3 = index.similarity_search(
query_text="Greek myths",
columns=["id", "field2"],
num_results=2,
query_type=" ? "
)Setting query_type to hybrid combines semantic ANN results with keyword results. ann is the vector-only default, and FULL_TEXT is keyword-only.
Source: docs.databricks.comFilters narrow results on indexed columns, and their syntax depends on the endpoint type. Standard endpoints take a filter dictionary, such as filters={"status": "active"}. Storage-optimized endpoints take a single SQL-like string, such as filters="status = 'active'". If you omit query_text, query_vector and query_type and pass only filters, the query returns matching rows with no retrieval at all. Filtering on the primary key this way gives a point lookup.
results = index.similarity_search(
query_text="Greek myths",
columns=["id", "text"],
filters="profile.age = 18 AND attrs['voltage'] = 9",
num_results=10
)Checkpoint 8 of 9· Check yourself
A high-throughput product search API has a strict latency budget. A teammate proposes adding the reranker to every query. What is the right assessment?
The reranker trades latency for precision. The docs describe it as potentially less suitable for high-throughput, low-latency search.
“The reranker typically improves quality by approximately 10% but adds latency.”Source: docs.databricks.com
Checkpoint 9 of 9· Exam question
An e-commerce team's product catalog table changes dozens of times per minute, and their RAG-based shopping assistant must reflect price and availability updates within seconds of a Delta table write. Which configuration should they choose?
Correct answer: A — Use a standard endpoint with a Delta Sync index in continuous sync mode, with Change Data Feed enabled on the source product catalog table
- A. Continuous sync on a standard endpoint maintains near real-time synchronization, on the order of seconds, which matches the requirement to reflect writes almost immediately. Change Data Feed is what allows the endpoint to detect and propagate those row-level changes as they happen.
- B. Storage-optimized endpoints only support triggered sync, and manually invoking a sync after dozens of updates per minute would either lag far behind the writes or require constant re-triggering, neither of which delivers second-level freshness reliably. This approach is built for infrequent, cost-sensitive updates rather than high-frequency ones.
- C. An hourly triggered sync leaves the index stale for up to an hour after any given catalog change, which is far outside the seconds-level freshness the shopping assistant needs. This cadence is suited to slowly changing data, not a fast-moving product catalog.
- D. Rebuilding a Direct Vector Access index from scratch on every table change is both operationally impractical at this update frequency and forgoes the incremental synchronization that a Delta Sync index provides automatically. It would introduce far more latency and load than continuous sync.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Vector Search ranks results by cosine similarity out of the box.Why is that wrong?
It uses L2 distance. To get cosine-equivalent rankings, you normalize the embeddings yourself before indexing.
Covered in What Vector Search is and how it measures similarity
2.You can start with self-managed embeddings and switch the same index to Databricks-managed embeddings later.Why is that wrong?
The conversion is not possible. Moving to managed embeddings means creating a new index and recomputing the embeddings.
3.A storage-optimized endpoint can keep an index in Continuous sync for second-level freshness.Why is that wrong?
Storage-optimized endpoints support only Triggered sync. Continuous sync is a standard-endpoint option.
4.A Direct Vector Access Index can be created from Catalog Explorer like a Delta Sync Index.Why is that wrong?
Direct Vector Access Indexes can only be created through the REST API or the SDK.
5.Because hybrid is the recommended starting point, similarity_search runs hybrid search by default.Why is that wrong?
The default is ANN. Hybrid search only runs when query_type is set to hybrid.
Covered in Querying: ANN, full-text, hybrid, reranking and filters
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Databricks AI Search (formerly Databricks Vector Search) is a search solution that is built into the Databricks Data + AI Platform”
↩︎ What Vector Search is and how it measures similarity“AI Search uses the Hierarchical Navigable Small World (HNSW) algorithm for its approximate nearest neighbor (ANN) searches and the L2 distance metric”
↩︎ What Vector Search is and how it measures similarity“The similarity search and keyword search results are combined using the Reciprocal Rank Fusion (RRF) function.”
↩︎ What Vector Search is and how it measures similarity“Relevance scores are calculated using Okapi BM25.”
↩︎ What Vector Search is and how it measures similarity“Endpoints scale up automatically to support the size of the index or the number of concurrent requests.”
↩︎ The components: endpoint, index and embedding model endpoint“AI Search indexes appear in and are governed by Unity Catalog.”
↩︎ The components: endpoint, index and embedding model endpoint“you can use a pre-configured Foundation Model APIs endpoint or create a model serving endpoint to serve the embedding model of your choice.”
↩︎ The components: endpoint, index and embedding model endpoint“Standard endpoints have a capacity of 320 million vectors at dimension 768.”
↩︎ The components: endpoint, index and embedding model endpoint“Storage-optimized endpoints have a larger capacity (over one billion vectors at dimension 768) and provide 10-20x faster indexing.”
↩︎ The components: endpoint, index and embedding model endpoint“To create an AI Search index, you must have CREATE TABLE privileges on the catalog schema where the index will be created.”
↩︎ The components: endpoint, index and embedding model endpoint“Databricks recommends that you use service principals, which can have a per-query performance up to 100 msec faster relative to personal access tokens”
↩︎ The components: endpoint, index and embedding model endpoint“Databricks calculates the embeddings, using a model that you specify, and optionally saves the embeddings to a table in Unity Catalog.”
↩︎ Index types, embedding sources and sync modes“Dedicated full-text search indexes are only available on storage-optimized endpoints and require triggered sync mode.”
↩︎ Index types, embedding sources and sync modes“An AI Search endpoint. This endpoint serves the AI Search index.”
↩︎ Key concept“If you want to use cosine similarity you need to normalize your datapoint embeddings before feeding them into the vector search algorithm.”
↩︎ Exam trap 1“It is not possible to convert a self-managed embedding index to a Databricks-managed index.”
↩︎ Exam trap 2“When the data points are normalized, the ranking produced by L2 distance is the same as the ranking produced by cosine similarity.”
↩︎ Prediction“useful in RAG applications where source data has unique keywords such as SKUs or identifiers that are not well suited to pure similarity search”
↩︎ Checkpoint“For standard endpoints, the source table must use a change data feed.”
↩︎ Checkpoint - 2.
“Delta Sync Index automatically syncs with a source table, incrementally updating the index as the underlying data in the source table changes.”
↩︎ Index types, embedding sources and sync modes“The user is responsible for updating this table using the REST API or the Python SDK.”
↩︎ Index types, embedding sources and sync modes“However, it has a higher cost associated with it since a compute cluster is provisioned to run the continuous sync streaming pipeline.”
↩︎ Index types, embedding sources and sync modes“The name of the table is the name of the AI Search index, appended by _writeback_table.”
↩︎ Index types, embedding sources and sync modes“The primary key and embedding columns are always included in the index.”
↩︎ Index types, embedding sources and sync modes“The index schema is fixed at creation time, so any schema changes require creating a new index to take effect.”
↩︎ Index types, embedding sources and sync modes“For storage-optimized endpoints, only Triggered sync mode is supported.”
↩︎ Exam trap 3“This type of index cannot be created using the UI. You must use the REST API or the SDK.”
↩︎ Exam trap 4“Direct Vector Access Index supports direct read and write of vectors and metadata.”
↩︎ Checkpoint“After the new index is ready, switch traffic to the new index.”
↩︎ Checkpoint - 3.
“You can only query the AI Search index using the Python SDK, the REST API, or the SQL vector_search() AI function.”
↩︎ Querying: ANN, full-text, hybrid, reranking and filters“General-purpose retrieval. The recommended starting point for most use cases.”
↩︎ Querying: ANN, full-text, hybrid, reranking and filters“Databricks recommends trying out reranking for any RAG agent use case.”
↩︎ Querying: ANN, full-text, hybrid, reranking and filters“adopt a more SQL-like filter string instead of the filter dictionary used in standard AI Search endpoints”
↩︎ Querying: ANN, full-text, hybrid, reranking and filters“The default query type is ann (approximate nearest neighbor).”
↩︎ Exam trap 5“The default query type is ann (approximate nearest neighbor).”
↩︎ Prediction“The reranker typically improves quality by approximately 10% but adds latency.”
↩︎ Checkpoint