What you will be able to do
- Name the three ways to query an AI Search index and the Unity Catalog privileges a non-owner needs
- Write similarity_search calls with query_text or query_vector, and choose between ANN, hybrid and FULL_TEXT
- Write filters for standard and storage-optimized endpoints and explain why a filtered query can come back empty
- Handle results beyond 1,000 rows with the SDK and the REST API
1.Who can query, and through which interface
An index on Databricks AI Search (formerly Vector Search) can be queried in exactly three ways: the Python SDK, the REST API, or the SQL vector_search() function, which is in Public Preview. In the SDK you call client.get_index(index_name=...) to get the index object and then similarity_search() on it. Over REST you call /api/2.0/vector-search/indexes/{index_name}/query.
Someone who does not own the index needs USE CATALOG, USE SCHEMA, and SELECT on the index. Two authentication methods are supported. A personal access token works, and in a notebook the SDK generates one for you. For production, Databricks recommends a service principal: it is more secure and can save up to 100 msec per query. The REST example gets an OAuth token scoped to the ReadVectorIndex operation on that index.
Checkpoint 1 of 6· Check yourself
Which of these is NOT a supported way to query an AI Search index?
The docs list only the SDK, the REST API and vector_search(). The index is not read like an ordinary table.
“You can only query the AI Search index using the Python SDK, the REST API, or the SQL vector_search() AI function.”Source: docs.databricks.com
Sources1
2.query_text vs query_vector, and ANN vs hybrid vs full-text
What you put in the query depends on how the index was built. If Databricks computes the embeddings, you send query_text and the configured model embeds it for you. If the index holds precomputed embeddings, you send query_vector, embedded with the same dimension as the index. columns chooses which fields come back, and num_results sets how many.
Checkpoint 2 of 6· Fill the gap
This index holds precomputed 1024-dimension embeddings. Which parameter completes the query?
results2 = index.similarity_search(
? =[0.9] * 1024,
columns=["id", "text"],
num_results=2
)When the index uses self-managed embeddings, you pass the query as a vector. query_text only works when Databricks computes the embeddings.
Source: docs.databricks.comquery_type controls how matches are found. The default, ANN, runs approximate nearest neighbour search with HNSW and L2 distance. If you want cosine similarity, normalize the embeddings first; normalized vectors rank the same under L2 and cosine. hybrid runs a vector search and a BM25 keyword search and combines the two with Reciprocal Rank Fusion. It is the right choice when the data contains exact identifiers such as SKUs. FULL_TEXT (Beta) matches keywords only, without embeddings.
# Delta Sync Index using hybrid search, with embeddings computed by Databricks
results3 = index.similarity_search(
query_text="Greek myths",
columns=["id", "field2"],
num_results=2,
query_type="hybrid"
)| query_type | How it matches | Result limit noted in docs |
|---|---|---|
| ann (default) | Approximate nearest neighbour over embeddings | 10,000 across all pages |
| hybrid | Vector search plus BM25 keyword search, fused with RRF | 200 by default |
| FULL_TEXT (Beta) | Keyword matching only, no embeddings | Up to 10,000 |
Checkpoint 3 of 6· Exam question
When configuring a Delta Sync index to have Databricks compute embeddings automatically (rather than supplying pre-computed vectors), what requirement applies to the `embedding_source_column`?
Correct answer: A — The column must contain text (string) data, since the embedding model endpoint generates vectors from raw text.
- A. Databricks-computed embeddings require a text column, since the specified embedding model endpoint reads the raw string content and produces the vector representation itself.
- B. An array of floats describes a pre-computed embedding, which is the self-managed embeddings path using `embedding_vector_column`, not the Databricks-computed path that reads text.
- C. There is no requirement to normalize the source column into a numeric range; the embedding model endpoint consumes raw text and handles vectorization internally.
- D. A struct combining text and a vector is not the expected schema; Databricks-computed embeddings only need a plain text column, and any pre-computed vector would be redundant here.
3.Filters: dictionary on standard, SQL string on storage-optimized
filters restricts results by the values in indexed columns. The syntax depends on the type of endpoint the index lives on. A standard endpoint takes a Python dict, where each key holds both the column and the operator, and keys are combined with AND. A storage-optimized endpoint takes a string written like a SQL WHERE clause.
# Standard endpoint syntax
results = index.similarity_search(
query_text="Greek myths",
columns=all_columns,
filters={"id NOT": ("13770", "88231")},
num_results=2)| Intent | Standard (dict) | Storage-optimized (string) |
|---|---|---|
| Exact match | {"make": "Toyota"} | make = 'Toyota' |
| Negation | {"make NOT": "Ford"} | make != 'Ford' |
| Any of several values | {"make": ["Toyota","Honda"]} | make IN ('Toyota','Honda') |
| Range | {"price >=": 30000, "price <=": 55000} | price >= 30000 AND price <= 55000 |
| LIKE | Whole tokens only: {"color LIKE": "red"} | Wildcards: color LIKE 'bl%' |
| Array contains | {"body_type": "sedan"} | Not supported |
The dictionary form has some pitfalls. There is no BETWEEN operator. A Python dict keeps only the last of two identical keys, so two LIKE conditions on the same column must go in a list of dicts. You also cannot call SQL functions inside the dict. The storage-optimized form has its own catch: it fetches extra candidates and applies the filter to them afterwards. If no matching row scores high enough to be among those candidates, the query returns nothing, even though matching rows exist.
Checkpoint 4 of 6· Check yourself
On a standard endpoint, filters={"title LIKE": "%2024%", "title LIKE": "%Tesla%"} returns rows that don't mention 2024. Why?
In a dict, a repeated key overwrites the earlier one. To apply both conditions to the same field, pass a list of dicts.
“Python keeps only the last value for duplicate keys in a dict.”Source: docs.databricks.com
Checkpoint 5 of 6· Exam question
A team is building a customer support agent where new support ticket resolutions must be searchable within a few seconds of being written to the source Delta table, and the team is willing to pay extra for lower latency. Which `pipeline_type` should they choose when creating the Delta Sync index?
Correct answer: A — CONTINUOUS, because it keeps the index synced with seconds of latency at a higher ongoing cost.
- A. Continuous pipelines keep the index synced with only seconds of latency, which matches the near-real-time requirement, at the cost of higher ongoing compute compared to triggered syncing.
- B. A triggered pipeline does not sync on its own; it requires an explicit `sync()` call, so it would not deliver the automatic few-second freshness the team needs.
- C. Continuous pipelines are designed for low-latency syncing, not daily batching; batching updates once per day describes a manual triggered workflow, not the continuous option.
- D. Storage-optimized endpoints actually only support the triggered pipeline type, not continuous, so this option describes the relationship backwards and would not meet the latency goal anyway.
Sources3
4.Getting more than 1,000 results
Any query type returns at most 1,000 results per page, and at most 10,000 across all pages. The SDK takes care of paging for you: set num_results=5000 and you get all 5,000 back. With the REST API you page yourself. Each response that has more results includes a next_page_token. You pass that token to /query-next-page and keep going until no token comes back.
Checkpoint 6 of 6· Check yourself
A REST client requests num_results=5000 and only gets 1,000 rows. What should it do?
Over REST, paging is manual and uses the token from each response. The SDK does the same thing for you behind the scenes.
“When you use the REST API directly, you must handle pagination manually.”Source: docs.databricks.com
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The dictionary filter you wrote for a standard endpoint will work unchanged on a storage-optimized endpoint.Why is that wrong?
The two endpoint types use different syntax: a dict on standard endpoints and a SQL-like string on storage-optimized endpoints.
Covered in Filters: dictionary on standard, SQL string on storage-optimized
2.If a filtered query on a storage-optimized endpoint returns nothing, no rows match the filter.Why is that wrong?
The filter is applied after an over-fetched candidate set is retrieved. Matching rows that score too low to be among the candidates never reach the filter.
Covered in Filters: dictionary on standard, SQL string on storage-optimized
Practise it for real
Build a Delta Sync Index with Databricks-computed embeddings over a small Wikipedia sample, then run a plain query and a filtered one.
1.Run %pip install databricks-ai-search, restart Python, then create vsc = AISearchClient().
Why: AISearchClient is the SDK's entry point for both endpoints and indexes.
You should see: help(AISearchClient) prints the client's methods.
2.Read 10 rows from the en_wikipedia parquet dataset and save them as a Delta table with .option("delta.enableChangeDataFeed", "true").
Why: Standard endpoints need a change data feed on the source table.
You should see: SELECT * on the table returns 10 articles with id and text columns.
3.Call vsc.create_endpoint(name=..., endpoint_type="STANDARD"), then vsc.get_endpoint(name=...).
Why: The index has to be attached to an existing endpoint.
You should see: get_endpoint returns the endpoint's details.
4.Call vsc.create_delta_sync_index with pipeline_type='TRIGGERED', primary_key="id", embedding_source_column="text" and embedding_model_endpoint_name="databricks-qwen3-embedding-0-6b", then poll describe().
Why: Databricks embeds the text column. Queries should wait until the index is online.
You should see: After several minutes the loop prints "Index is ONLINE".
5.Run index.similarity_search(query_text="Greek myths", columns=all_columns, num_results=2), then run it again with filters={"id NOT": ("13770", "88231")}.
Why: This exercises a basic ANN query and the standard-endpoint dict filter syntax.
You should see: Each call returns a result with manifest columns and a data_array of rows. The filtered call leaves out those two ids.
Stuck? Get a nudge
If the first query times out, check whether the embedding endpoint scaled to zero.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“USE CATALOG on the catalog that contains the AI Search index.”
↩︎ Who can query, and through which interface“using service principals can improve performance by up to 100 msec per query.”
↩︎ Who can query, and through which interface“The vector_search() AI function is in Public Preview.”
↩︎ Who can query, and through which interface“The default query type is ann (approximate nearest neighbor).”
↩︎ query_text vs query_vector, and ANN vs hybrid vs full-text“By default, hybrid search includes all text metadata columns and returns a maximum of 200 results.”
↩︎ query_text vs query_vector, and ANN vs hybrid vs full-text“The maximum number of results that a single query can return across all pages is 10,000.”
↩︎ Getting more than 1,000 results“The Python SDK handles pagination transparently.”
↩︎ Getting more than 1,000 results“It is possible that no results will be returned even if there are results in the dataset that match the filter condition”
↩︎ Exam trap 2“SELECT on the AI Search index.”
↩︎ Prediction“You can only query the AI Search index using the Python SDK, the REST API, or the SQL vector_search() AI function.”
↩︎ Checkpoint“When you use the REST API directly, you must handle pagination manually.”
↩︎ Checkpoint - 2.
“If you want to use cosine similarity you need to normalize your datapoint embeddings before feeding them into the vector search algorithm.”
↩︎ query_text vs query_vector, and ANN vs hybrid vs full-text“This method is particularly useful in RAG applications where source data has unique keywords such as SKUs or identifiers”
↩︎ query_text vs query_vector, and ANN vs hybrid vs full-text - 3.
“Multiple keys in a single dict are combined with AND logic.”
↩︎ Filters: dictionary on standard, SQL string on storage-optimized“Array filtering is not supported on storage-optimized endpoints.”
↩︎ Filters: dictionary on standard, SQL string on storage-optimized“There is no BETWEEN operator.”
↩︎ Filters: dictionary on standard, SQL string on storage-optimized“standard endpoints use a Python dictionary, and storage-optimized endpoints use a SQL-like filter string.”
↩︎ Exam trap 1“Python keeps only the last value for duplicate keys in a dict.”
↩︎ Checkpoint