CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 4 · Lesson 35/56

    Mosaic AI Vector Search: Endpoints, Indexes, Embeddings and Query Types

    Explain the key concepts and components of Mosaic AI Vector Search

    17 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Name the two components you must create to use Vector Search (an endpoint and an index), the optional embedding model endpoint, and how they relate.
    • Tell a Delta Sync Index from a Direct Vector Access Index, and pick between Databricks-managed embeddings, self-managed embeddings and a full-text-only index.
    • Explain how similarity is computed (HNSW, L2 distance, normalization for cosine) and how hybrid search combines ANN and BM25 with Reciprocal Rank Fusion.
    • Choose between the ANN, full-text, hybrid and reranked query types, and know the default query type and the filter syntax for each endpoint type.

    Key concept

    Endpoint serves, index holds — Vector Search is built from two objects you create. An index holds the embedded rows of a source table and is governed in Unity Catalog. An endpoint is the serving compute that hosts and answers queries against one or more indexes.

    1.What Vector Search is and how it measures similarity

    The exam guide calls this product Mosaic AI Vector Search. The current Databricks documentation calls it Databricks AI Search and notes that it was formerly Databricks Vector Search. Treat these as names for the same product. The SDK and docs use AISearchClient, but the REST paths still read /api/2.0/vector-search/..., so expect both names on the exam.

    Vector Search is the retrieval layer for AI applications such as RAG systems, recommender systems and image recognition. It works on embeddings, which are numeric representations of the meaning of text or images, produced by a model. A query is turned into an embedding (or arrives as one), and the service returns the stored entries whose embeddings are closest to it, along with the documents attached to them. Because it is built into the platform, it shares the platform's governance: indexes live in Unity Catalog, and endpoints are controlled with access control lists.

    Search is approximate, not exhaustive. Vector Search uses the Hierarchical Navigable Small World (HNSW) algorithm for approximate nearest neighbor (ANN) search, and L2 distance as its similarity metric. The cosine point from the predict above is the detail most likely to be tested. Normalization does not happen automatically: if your design depends on cosine similarity, you normalize the vectors before they reach the index.

    Pure vector search has a known weakness. Embeddings capture meaning well but can blur exact tokens such as part numbers or error codes. So Vector Search also supports hybrid keyword-similarity search. A keyword search scores documents with Okapi BM25 across text and string columns, an ANN search runs alongside it, and the two ranked lists are merged with Reciprocal Rank Fusion (RRF). RRF rescores each document by its rank in each list (with rrf_param set to 60), normalizes the result so the best possible score is 1, and returns the documents with the highest combined scores.

    Checkpoint 1 of 9· Check yourself

    A RAG assistant for a parts catalogue often misses documents when users type exact SKUs, but handles conceptual questions well. Which Vector Search capability targets this gap?

    Sources1

    2.The components: endpoint, index and embedding model endpoint

    Using Vector Search means creating two things, plus an optional third.

    1. A Vector Search endpoint. This is the compute that serves indexes. You can manage it through the UI (Compute → AI Search tab), the Python SDK or the REST API. It scales up automatically with index size and with the number of concurrent requests, and scales down when an index is deleted. 2. A Vector Search index. This is the searchable structure built from a source table: a Delta Lake table, a streaming table or a managed Iceberg v3+ table. It stores the embedded data together with metadata, supports real-time ANN queries, and is a Unity Catalog object. Its name therefore uses the three-level <catalog>.<schema>.<name> form. 3. An embedding model serving endpoint (optional). You only need this if Databricks computes the embeddings for you. It can be a pre-configured Foundation Model APIs endpoint or a model serving endpoint that you create for an embedding model of your choice.

    You choose the endpoint type when you create it, and that choice limits what its indexes can do later:

    The two endpoint types and the property that defines each
    Endpoint typeCapacityIndexing and latencyConstraints
    STANDARD320 million vectors at dimension 768Standard query latency; high QPS available for sustained throughputSource table must use a change data feed
    STORAGE_OPTIMIZEDOver one billion vectors at dimension 76810-20x faster indexing; queries about 250 msec slowerTriggered sync only; filters use a SQL-like string
    Creating an endpoint with the SDK. With no arguments the client generates a PAT; the commented line shows service principal authentication.python
    # The following line automatically generates a PAT Token for authentication
    client = AISearchClient()
    
    # The following line uses the service principal token for authentication
    # client = AISearchClient(service_principal_client_id=<CLIENT_ID>,service_principal_client_secret=<CLIENT_SECRET>)
    
    client.create_endpoint(
        name="vector_search_endpoint_name",
        endpoint_type="STANDARD" # or "STORAGE_OPTIMIZED"
    )

    The surrounding requirements are part of the component picture:

    - Workspace: Unity Catalog must be enabled, and so must serverless compute. - Source table: for standard endpoints it needs a change data feed. Tables with row tracking enabled get one automatically. - Permissions: creating an index requires CREATE TABLE on the target schema. Creating and managing endpoints is controlled through endpoint ACLs. - Authentication: two modes are supported, service principals and personal access tokens (PATs). In a notebook, the SDK generates a PAT for you. For production, Databricks recommends service principals, which can also be up to 100 msec faster per query. - Security: data is encrypted at rest (AES-256) and in transit (TLS 1.2+).

    Checkpoint 2 of 9· Check yourself

    An engineer creates a Delta Sync Index on a standard endpoint. The source Delta table has neither a change data feed nor row tracking. What must change?

    Checkpoint 3 of 9· Exam question

    A data engineering team is building a RAG index over an internal knowledge base that is refreshed only twice a day by a batch ETL job. Cost efficiency matters more than index freshness, and the team does not need a continuously running sync process. Which AI Search endpoint and sync configuration best fits these requirements?

    Sources1

    3.Index types, embedding sources and sync modes

    Before creating an index you make two independent decisions: how the index stays current (the index type) and where the embeddings come from.

    Index type. - A Delta Sync Index stays tied to a source table and updates incrementally as that table changes. - A Direct Vector Access Index has no source table. You write vectors and metadata into it yourself with upsert and delete calls through the SDK or REST API. This type cannot be created in the UI.

    Embedding source. - Databricks-managed embeddings: the source holds text. Databricks computes embeddings with the model you name, and computes the query embedding at search time too. Optionally, it saves the computed embeddings to a Unity Catalog table named after the index with _writeback_table appended. That table is read-only and is deleted along with the index. - Self-managed embeddings: the source already holds an array<float> column of vectors, and you give its dimension. Queries then pass query_vector instead of query_text. - No embeddings (Beta): a dedicated full-text index for BM25 keyword search only, available only on storage-optimized endpoints.

    The embedding source is permanent. You cannot convert a self-managed index into a managed one; you would have to create a new index and recompute the embeddings.

    A Delta Sync Index with self-managed embeddings: you name the vector column and its dimension instead of an embedding model endpoint.python
    index = client.create_delta_sync_index(
      endpoint_name="vector_search_demo_endpoint",
      source_table_name="vector_search_demo.vector_search.en_wiki",
      index_name="vector_search_demo.vector_search.en_wiki_index",
      pipeline_type="TRIGGERED",
      primary_key="id",
      embedding_dimension=1024,
      embedding_vector_column="text_vector"
    )

    pipeline_type sets the sync mode of a Delta Sync Index:

    - Continuous keeps the index within seconds of the source table. It costs more because a compute cluster runs a streaming sync pipeline. - Triggered syncs only when you call index.sync(), click Sync now in Catalog Explorer, or call the REST sync endpoint.

    On standard endpoints both modes update incrementally. On storage-optimized endpoints, Triggered is the only mode, each sync partially rebuilds the index, and embeddings for unchanged rows are reused.

    A few structural rules come with the index:

    - Every index has a primary key. - columns_to_sync limits which source columns are indexed. The primary key and embedding column are always included, and only indexed columns can be returned or filtered on. - The column name _id is reserved. - The schema is fixed at creation. Adding or modifying source columns means building a new index.

    Checkpoint 4 of 9· Match them up

    Match each Vector Search component to its behaviour

    Tap a term, then the definition that fits it.

    Checkpoint 5 of 9· Put it in order

    Put the zero-downtime index rebuild steps in order

    1. 1.Delete the original index
    2. 2.Perform the schema change on your source table
    3. 3.After the new index is ready, switch traffic to the new index
    4. 4.Create a new index using the updated schema

    Checkpoint 6 of 9· Exam question

    A GenAI engineering team has already computed embeddings for their documents using a proprietary fine-tuned embedding model and stores the resulting vectors as a column in a Delta table alongside the document text. They want the index to automatically pick up new rows as the Delta table changes, without Databricks recomputing embeddings. Which index configuration should they use?

    Sources21

    4.Querying: ANN, full-text, hybrid, reranking and filters

    There are three ways to query an index: the Python SDK (similarity_search), the REST API, and the SQL vector_search() function. A user who does not own the index needs USE CATALOG, USE SCHEMA and SELECT on it. The query_type parameter selects the retrieval algorithm.

    Retrieval strategies selected with query_type, and what each is best for
    StrategyHow it worksBest for
    ANN (default)Searches using vector embeddings to find semantically similar documentsConceptual queries where meaning matters more than exact wording
    Full-text (FULL_TEXT, Beta)Keyword search that matches on exact termsProper nouns, product IDs, error codes, technical jargon
    HybridCombines ANN and full-text results using Reciprocal Rank Fusion (RRF)General-purpose retrieval; the recommended starting point
    Hybrid + rerankerRuns hybrid search, then re-scores results with a cross-encoder reranker modelHigher precision when latency allows

    Each strategy has limits worth knowing:

    - Hybrid searches all text metadata columns by default and returns at most 200 results. - Full-text can return up to 10,000 keyword matches without using embeddings, on both endpoint types. - Reranker: an optional second pass that can sit on top of any strategy. A cross-encoder re-scores each retrieved result against the query. It typically improves quality by about 10% at the cost of extra latency. That suits RAG chatbots, where Databricks recommends trying it, more than high-throughput, low-latency search.

    Which arguments a query passes depends on how the index was built. A managed-embedding index takes query_text, and a self-managed one takes query_vector.

    Checkpoint 7 of 9· Fill the gap

    Which query_type value runs ANN and keyword search in parallel and fuses them with RRF?

    results3 = index.similarity_search(
        query_text="Greek myths",
        columns=["id", "field2"],
        num_results=2,
        query_type=" ? "
        )

    Filters narrow results on indexed columns, and their syntax depends on the endpoint type. Standard endpoints take a filter dictionary, such as filters={"status": "active"}. Storage-optimized endpoints take a single SQL-like string, such as filters="status = 'active'". If you omit query_text, query_vector and query_type and pass only filters, the query returns matching rows with no retrieval at all. Filtering on the primary key this way gives a point lookup.

    On a storage-optimized endpoint, filters are a single SQL WHERE-style string rather than a dictionary.python
    results = index.similarity_search(
        query_text="Greek myths",
        columns=["id", "text"],
        filters="profile.age = 18 AND attrs['voltage'] = 9",
        num_results=10
    )

    Checkpoint 8 of 9· Check yourself

    A high-throughput product search API has a strict latency budget. A teammate proposes adding the reranker to every query. What is the right assessment?

    Checkpoint 9 of 9· Exam question

    An e-commerce team's product catalog table changes dozens of times per minute, and their RAG-based shopping assistant must reflect price and availability updates within seconds of a Delta table write. Which configuration should they choose?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Vector Search ranks results by cosine similarity out of the box.Why is that wrong?

      It uses L2 distance. To get cosine-equivalent rankings, you normalize the embeddings yourself before indexing.

      Covered in What Vector Search is and how it measures similarity

    2. 2.You can start with self-managed embeddings and switch the same index to Databricks-managed embeddings later.Why is that wrong?

      The conversion is not possible. Moving to managed embeddings means creating a new index and recomputing the embeddings.

      Covered in Index types, embedding sources and sync modes

    3. 3.A storage-optimized endpoint can keep an index in Continuous sync for second-level freshness.Why is that wrong?

      Storage-optimized endpoints support only Triggered sync. Continuous sync is a standard-endpoint option.

      Covered in Index types, embedding sources and sync modes

    4. 4.A Direct Vector Access Index can be created from Catalog Explorer like a Delta Sync Index.Why is that wrong?

      Direct Vector Access Indexes can only be created through the REST API or the SDK.

      Covered in Index types, embedding sources and sync modes

    5. 5.Because hybrid is the recommended starting point, similarity_search runs hybrid search by default.Why is that wrong?

      The default is ANN. Hybrid search only runs when query_type is set to hybrid.

      Covered in Querying: ANN, full-text, hybrid, reranking and filters

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Databricks AI Search (formerly Databricks Vector Search) is a search solution that is built into the Databricks Data + AI Platform”
      ↩︎ What Vector Search is and how it measures similarity
      “AI Search uses the Hierarchical Navigable Small World (HNSW) algorithm for its approximate nearest neighbor (ANN) searches and the L2 distance metric”
      ↩︎ What Vector Search is and how it measures similarity
      “The similarity search and keyword search results are combined using the Reciprocal Rank Fusion (RRF) function.”
      ↩︎ What Vector Search is and how it measures similarity
      “Relevance scores are calculated using Okapi BM25.”
      ↩︎ What Vector Search is and how it measures similarity
      “Endpoints scale up automatically to support the size of the index or the number of concurrent requests.”
      ↩︎ The components: endpoint, index and embedding model endpoint
      “AI Search indexes appear in and are governed by Unity Catalog.”
      ↩︎ The components: endpoint, index and embedding model endpoint
      “you can use a pre-configured Foundation Model APIs endpoint or create a model serving endpoint to serve the embedding model of your choice.”
      ↩︎ The components: endpoint, index and embedding model endpoint
      “Standard endpoints have a capacity of 320 million vectors at dimension 768.”
      ↩︎ The components: endpoint, index and embedding model endpoint
      “Storage-optimized endpoints have a larger capacity (over one billion vectors at dimension 768) and provide 10-20x faster indexing.”
      ↩︎ The components: endpoint, index and embedding model endpoint
      “To create an AI Search index, you must have CREATE TABLE privileges on the catalog schema where the index will be created.”
      ↩︎ The components: endpoint, index and embedding model endpoint
      “Databricks recommends that you use service principals, which can have a per-query performance up to 100 msec faster relative to personal access tokens”
      ↩︎ The components: endpoint, index and embedding model endpoint
      “Databricks calculates the embeddings, using a model that you specify, and optionally saves the embeddings to a table in Unity Catalog.”
      ↩︎ Index types, embedding sources and sync modes
      “Dedicated full-text search indexes are only available on storage-optimized endpoints and require triggered sync mode.”
      ↩︎ Index types, embedding sources and sync modes
      “An AI Search endpoint. This endpoint serves the AI Search index.”
      ↩︎ Key concept
      “If you want to use cosine similarity you need to normalize your datapoint embeddings before feeding them into the vector search algorithm.”
      ↩︎ Exam trap 1
      “It is not possible to convert a self-managed embedding index to a Databricks-managed index.”
      ↩︎ Exam trap 2
      “When the data points are normalized, the ranking produced by L2 distance is the same as the ranking produced by cosine similarity.”
      ↩︎ Prediction
      “useful in RAG applications where source data has unique keywords such as SKUs or identifiers that are not well suited to pure similarity search”
      ↩︎ Checkpoint
      “For standard endpoints, the source table must use a change data feed.”
      ↩︎ Checkpoint
    2. 2.
      “Delta Sync Index automatically syncs with a source table, incrementally updating the index as the underlying data in the source table changes.”
      ↩︎ Index types, embedding sources and sync modes
      “The user is responsible for updating this table using the REST API or the Python SDK.”
      ↩︎ Index types, embedding sources and sync modes
      “However, it has a higher cost associated with it since a compute cluster is provisioned to run the continuous sync streaming pipeline.”
      ↩︎ Index types, embedding sources and sync modes
      “The name of the table is the name of the AI Search index, appended by _writeback_table.”
      ↩︎ Index types, embedding sources and sync modes
      “The primary key and embedding columns are always included in the index.”
      ↩︎ Index types, embedding sources and sync modes
      “The index schema is fixed at creation time, so any schema changes require creating a new index to take effect.”
      ↩︎ Index types, embedding sources and sync modes
      “For storage-optimized endpoints, only Triggered sync mode is supported.”
      ↩︎ Exam trap 3
      “This type of index cannot be created using the UI. You must use the REST API or the SDK.”
      ↩︎ Exam trap 4
      “Direct Vector Access Index supports direct read and write of vectors and metadata.”
      ↩︎ Checkpoint
      “After the new index is ready, switch traffic to the new index.”
      ↩︎ Checkpoint
    3. 3.
      “You can only query the AI Search index using the Python SDK, the REST API, or the SQL vector_search() AI function.”
      ↩︎ Querying: ANN, full-text, hybrid, reranking and filters
      “General-purpose retrieval. The recommended starting point for most use cases.”
      ↩︎ Querying: ANN, full-text, hybrid, reranking and filters
      “Databricks recommends trying out reranking for any RAG agent use case.”
      ↩︎ Querying: ANN, full-text, hybrid, reranking and filters
      “adopt a more SQL-like filter string instead of the filter dictionary used in standard AI Search endpoints”
      ↩︎ Querying: ANN, full-text, hybrid, reranking and filters
      “The default query type is ann (approximate nearest neighbor).”
      ↩︎ Exam trap 5
      “The default query type is ann (approximate nearest neighbor).”
      ↩︎ Prediction
      “The reranker typically improves quality by approximately 10% but adds latency.”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.