CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 4 · Lesson 33/56

    Query a Vector Search (AI Search) Index: ANN, Hybrid, Filters and Pagination

    Create and query a Vector Search index

    10 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Name the three ways to query an AI Search index and the Unity Catalog privileges a non-owner needs
    • Write similarity_search calls with query_text or query_vector, and choose between ANN, hybrid and FULL_TEXT
    • Write filters for standard and storage-optimized endpoints and explain why a filtered query can come back empty
    • Handle results beyond 1,000 rows with the SDK and the REST API

    1.Who can query, and through which interface

    An index on Databricks AI Search (formerly Vector Search) can be queried in exactly three ways: the Python SDK, the REST API, or the SQL vector_search() function, which is in Public Preview. In the SDK you call client.get_index(index_name=...) to get the index object and then similarity_search() on it. Over REST you call /api/2.0/vector-search/indexes/{index_name}/query.

    Someone who does not own the index needs USE CATALOG, USE SCHEMA, and SELECT on the index. Two authentication methods are supported. A personal access token works, and in a notebook the SDK generates one for you. For production, Databricks recommends a service principal: it is more secure and can save up to 100 msec per query. The REST example gets an OAuth token scoped to the ReadVectorIndex operation on that index.

    Checkpoint 1 of 6· Check yourself

    Which of these is NOT a supported way to query an AI Search index?

    Sources1

    2.query_text vs query_vector, and ANN vs hybrid vs full-text

    What you put in the query depends on how the index was built. If Databricks computes the embeddings, you send query_text and the configured model embeds it for you. If the index holds precomputed embeddings, you send query_vector, embedded with the same dimension as the index. columns chooses which fields come back, and num_results sets how many.

    Checkpoint 2 of 6· Fill the gap

    This index holds precomputed 1024-dimension embeddings. Which parameter completes the query?

    results2 = index.similarity_search(
         ? =[0.9] * 1024,
        columns=["id", "text"],
        num_results=2
        )

    query_type controls how matches are found. The default, ANN, runs approximate nearest neighbour search with HNSW and L2 distance. If you want cosine similarity, normalize the embeddings first; normalized vectors rank the same under L2 and cosine. hybrid runs a vector search and a BM25 keyword search and combines the two with Reciprocal Rank Fusion. It is the right choice when the data contains exact identifiers such as SKUs. FULL_TEXT (Beta) matches keywords only, without embeddings.

    A hybrid keyword-plus-vector query on a Delta Sync Indexpython
    # Delta Sync Index using hybrid search, with embeddings computed by Databricks
    results3 = index.similarity_search(
        query_text="Greek myths",
        columns=["id", "field2"],
        num_results=2,
        query_type="hybrid"
        )
    Query types and their result limits
    query_typeHow it matchesResult limit noted in docs
    ann (default)Approximate nearest neighbour over embeddings10,000 across all pages
    hybridVector search plus BM25 keyword search, fused with RRF200 by default
    FULL_TEXT (Beta)Keyword matching only, no embeddingsUp to 10,000

    Checkpoint 3 of 6· Exam question

    When configuring a Delta Sync index to have Databricks compute embeddings automatically (rather than supplying pre-computed vectors), what requirement applies to the `embedding_source_column`?

    Sources12

    3.Filters: dictionary on standard, SQL string on storage-optimized

    filters restricts results by the values in indexed columns. The syntax depends on the type of endpoint the index lives on. A standard endpoint takes a Python dict, where each key holds both the column and the operator, and keys are combined with AND. A storage-optimized endpoint takes a string written like a SQL WHERE clause.

    Standard endpoint: dictionary filter that excludes two idspython
    # Standard endpoint syntax
    results = index.similarity_search(
      query_text="Greek myths",
      columns=all_columns,
      filters={"id NOT": ("13770", "88231")},
      num_results=2)
    The same filter intent in each endpoint's syntax
    IntentStandard (dict)Storage-optimized (string)
    Exact match{"make": "Toyota"}make = 'Toyota'
    Negation{"make NOT": "Ford"}make != 'Ford'
    Any of several values{"make": ["Toyota","Honda"]}make IN ('Toyota','Honda')
    Range{"price >=": 30000, "price <=": 55000}price >= 30000 AND price <= 55000
    LIKEWhole tokens only: {"color LIKE": "red"}Wildcards: color LIKE 'bl%'
    Array contains{"body_type": "sedan"}Not supported

    The dictionary form has some pitfalls. There is no BETWEEN operator. A Python dict keeps only the last of two identical keys, so two LIKE conditions on the same column must go in a list of dicts. You also cannot call SQL functions inside the dict. The storage-optimized form has its own catch: it fetches extra candidates and applies the filter to them afterwards. If no matching row scores high enough to be among those candidates, the query returns nothing, even though matching rows exist.

    Checkpoint 4 of 6· Check yourself

    On a standard endpoint, filters={"title LIKE": "%2024%", "title LIKE": "%Tesla%"} returns rows that don't mention 2024. Why?

    Checkpoint 5 of 6· Exam question

    A team is building a customer support agent where new support ticket resolutions must be searchable within a few seconds of being written to the source Delta table, and the team is willing to pay extra for lower latency. Which `pipeline_type` should they choose when creating the Delta Sync index?

    Sources3

    4.Getting more than 1,000 results

    Any query type returns at most 1,000 results per page, and at most 10,000 across all pages. The SDK takes care of paging for you: set num_results=5000 and you get all 5,000 back. With the REST API you page yourself. Each response that has more results includes a next_page_token. You pass that token to /query-next-page and keep going until no token comes back.

    Checkpoint 6 of 6· Check yourself

    A REST client requests num_results=5000 and only gets 1,000 rows. What should it do?

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The dictionary filter you wrote for a standard endpoint will work unchanged on a storage-optimized endpoint.Why is that wrong?

      The two endpoint types use different syntax: a dict on standard endpoints and a SQL-like string on storage-optimized endpoints.

      Covered in Filters: dictionary on standard, SQL string on storage-optimized

    2. 2.If a filtered query on a storage-optimized endpoint returns nothing, no rows match the filter.Why is that wrong?

      The filter is applied after an over-fetched candidate set is retrieved. Matching rows that score too low to be among the candidates never reach the filter.

      Covered in Filters: dictionary on standard, SQL string on storage-optimized

    Practise it for real

    Build a Delta Sync Index with Databricks-computed embeddings over a small Wikipedia sample, then run a plain query and a filtered one.

    1. 1.Run %pip install databricks-ai-search, restart Python, then create vsc = AISearchClient().

      Why: AISearchClient is the SDK's entry point for both endpoints and indexes.

      You should see: help(AISearchClient) prints the client's methods.

    2. 2.Read 10 rows from the en_wikipedia parquet dataset and save them as a Delta table with .option("delta.enableChangeDataFeed", "true").

      Why: Standard endpoints need a change data feed on the source table.

      You should see: SELECT * on the table returns 10 articles with id and text columns.

    3. 3.Call vsc.create_endpoint(name=..., endpoint_type="STANDARD"), then vsc.get_endpoint(name=...).

      Why: The index has to be attached to an existing endpoint.

      You should see: get_endpoint returns the endpoint's details.

    4. 4.Call vsc.create_delta_sync_index with pipeline_type='TRIGGERED', primary_key="id", embedding_source_column="text" and embedding_model_endpoint_name="databricks-qwen3-embedding-0-6b", then poll describe().

      Why: Databricks embeds the text column. Queries should wait until the index is online.

      You should see: After several minutes the loop prints "Index is ONLINE".

    5. 5.Run index.similarity_search(query_text="Greek myths", columns=all_columns, num_results=2), then run it again with filters={"id NOT": ("13770", "88231")}.

      Why: This exercises a basic ANN query and the standard-endpoint dict filter syntax.

      You should see: Each call returns a result with manifest columns and a data_array of rows. The filtered call leaves out those two ids.

    Stuck? Get a nudge

    If the first query times out, check whether the embedding endpoint scaled to zero.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “USE CATALOG on the catalog that contains the AI Search index.”
      ↩︎ Who can query, and through which interface
      “using service principals can improve performance by up to 100 msec per query.”
      ↩︎ Who can query, and through which interface
      “The vector_search() AI function is in Public Preview.”
      ↩︎ Who can query, and through which interface
      “The default query type is ann (approximate nearest neighbor).”
      ↩︎ query_text vs query_vector, and ANN vs hybrid vs full-text
      “By default, hybrid search includes all text metadata columns and returns a maximum of 200 results.”
      ↩︎ query_text vs query_vector, and ANN vs hybrid vs full-text
      “The maximum number of results that a single query can return across all pages is 10,000.”
      ↩︎ Getting more than 1,000 results
      “The Python SDK handles pagination transparently.”
      ↩︎ Getting more than 1,000 results
      “It is possible that no results will be returned even if there are results in the dataset that match the filter condition”
      ↩︎ Exam trap 2
      “SELECT on the AI Search index.”
      ↩︎ Prediction
      “You can only query the AI Search index using the Python SDK, the REST API, or the SQL vector_search() AI function.”
      ↩︎ Checkpoint
      “When you use the REST API directly, you must handle pagination manually.”
      ↩︎ Checkpoint
    2. 2.
      “If you want to use cosine similarity you need to normalize your datapoint embeddings before feeding them into the vector search algorithm.”
      ↩︎ query_text vs query_vector, and ANN vs hybrid vs full-text
      “This method is particularly useful in RAG applications where source data has unique keywords such as SKUs or identifiers”
      ↩︎ query_text vs query_vector, and ANN vs hybrid vs full-text
    3. 3.
      “Multiple keys in a single dict are combined with AND logic.”
      ↩︎ Filters: dictionary on standard, SQL string on storage-optimized
      “Array filtering is not supported on storage-optimized endpoints.”
      ↩︎ Filters: dictionary on standard, SQL string on storage-optimized
      “standard endpoints use a Python dictionary, and storage-optimized endpoints use a SQL-like filter string.”
      ↩︎ Exam trap 1
      “Python keeps only the last value for duplicate keys in a dict.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 6 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.