CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 4 · Lesson 37/56

    Vector Search Endpoint Sizing: Standard vs Storage-Optimized by Scale, Latency and Cost

    Configure vector search for a particular solution based on number of embeddings, update frequency, latency, and cost requirements.

    8 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Estimate how many vector search units an index needs from its vector count and dimension
    • Choose between standard and storage-optimized endpoints from vector count, latency target and cost per vector
    • Explain how embedding dimensionality affects latency and QPS
    • Describe how AI Search endpoints are billed and how to keep the base cost down

    Key concept

    Vector search unit (VSU) — The unit of capacity that sizes and bills an AI Search endpoint. A standard VSU holds about 2M vectors at dimension 768, and a storage-optimized VSU holds about 64M. An endpoint grows by whole units as its indexes grow, so vector count and dimension together decide capacity, performance and cost.

    1.Start from the number of embeddings

    Databricks AI Search was formerly called Databricks Vector Search, and the exam may use either name. When you configure it for a solution, the first thing to settle is the endpoint type, and you pick it when you create the endpoint. There are two: Standard and Storage Optimized. The main input to that choice is how many embeddings you need to store.

    Each type is measured in vector search units (VSUs). One standard VSU holds up to 2 million vectors at dimension 768. One storage-optimized VSU holds up to 64 million vectors at dimension 768. The phrase "or the equivalent" matters: what counts is total vector data, not just the row count.

    Endpoint capacity sets a hard limit on the choice. A standard endpoint holds 320 million vectors at dimension 768. A storage-optimized endpoint holds over one billion and indexes 10–20x faster. Both endpoint types have a base price and scale up automatically to match the total size of the indexes they serve. They scale back down when an index is deleted, and the smallest possible endpoint is one VSU.

    You set the type at creation, so make the decision before you build anything:

    The endpoint type is set when the endpoint is createdpython
    client.create_endpoint(
        name="vector_search_endpoint_name",
        endpoint_type="STANDARD" # or "STORAGE_OPTIMIZED"
    )

    Checkpoint 1 of 5· Check yourself

    A team expects about 600 million embeddings of dimension 768. Which endpoint type can hold them?

    Sources1

    2.Trade latency against cost per vector

    Capacity is only part of the decision. The two types also have very different latency. The performance guide states the rule directly: choose standard when latency is critical and the index is well under 320M vectors. Choose storage-optimized when you have 10M+ vectors, can accept extra latency, and need better cost efficiency per vector, which can be up to 7x cheaper. The reference numbers below were measured with self-managed embeddings at dimension 768.

    Reference performance by endpoint type (dimension 768)
    SKUVectorsLatencyQPS
    Standard10K20ms200+
    Standard10M40ms30
    Standard100M50ms30
    Storage-optimized10M300ms50
    Storage-optimized100M400ms40
    Storage-optimized1B500ms30

    Notice what happens as standard indexes grow. At 10K vectors, QPS is 200+, but by 10M it has fallen to 30. Performance is best when the index fits inside one VSU with room to spare. Once it grows past one unit (2M+ vectors on standard, 64M+ on storage-optimized), latency rises and QPS levels off at around 30 for ANN queries. So vector count shapes performance even within one endpoint type.

    Checkpoint 2 of 5· Match them up

    Match each endpoint characteristic to its value

    Tap a term, then the definition that fits it.

    Checkpoint 3 of 5· Exam question

    A team is building a RAG chatbot over a product catalog that changes every few seconds. Business stakeholders require the index to reflect catalog updates within seconds and are willing to accept higher compute cost for this freshness. Which vector search configuration best meets this requirement?

    Sources2

    3.Dimension is a sizing lever too

    Embedding dimensions are usually 384, 768, 1024 or 1536. Higher dimensions can capture more meaning and improve retrieval quality, but each query costs more compute. Dimension also affects the count side of the decision, because a VSU is defined at dimension 768 or its equivalent. Larger vectors fill units faster.

    The guide's rule is to use the smallest dimension that keeps retrieval quality acceptable. That makes 384 dimensions and below especially attractive for high-throughput, latency-sensitive applications. One constraint is specific to storage-optimized endpoints: if you supply precomputed embeddings, the dimension must be evenly divisible by 16.

    Checkpoint 4 of 5· Check yourself

    A latency-sensitive RAG app gets the same retrieval quality from a 384-dimension model as from a 1536-dimension model. What does the performance guide recommend?

    Sources23

    4.How endpoints are billed, and how to pay less

    AI Search bills for two things: the indexes that store your vectors and the endpoints that serve queries. An endpoint has a base price and then grows in whole VSUs. Billing starts only once an index exists on the endpoint. After the last index is deleted, charges stop 24 hours later.

    Because every endpoint carries its own base price, consolidation is the main cost lever. One endpoint can serve up to 50 indexes. If you expect low QPS across all of them, putting them on one endpoint means you pay one base price instead of several. The reverse also applies: endpoints whose indexes get no query traffic still incur serving costs, so find them and remove them. You can track spending with the system.billing.usage table (filtered to billing_origin_product = 'VECTOR_SEARCH'), the usage dashboards, and usage policies that tag endpoint and index costs by team or project.

    Checkpoint 5 of 5· Check yourself

    An engineer creates a standard endpoint on Monday and adds the first index on Thursday. When does the endpoint start being charged?

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.An AI Search endpoint starts billing as soon as it is created, and an idle endpoint costs nothing.Why is that wrong?

      Billing starts only after an index is created on the endpoint. From then on, an endpoint whose indexes get no queries still incurs serving costs until the indexes are removed.

      Covered in How endpoints are billed, and how to pay less

    2. 2.Higher-dimensional embeddings are always the better configuration.Why is that wrong?

      Higher dimensions add compute per query and fill VSUs faster. Use the smallest dimension that keeps retrieval quality acceptable.

      Covered in Dimension is a sizing lever too

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “each endpoint has a base price and scales up automatically to match the total size of the indices it is serving.”
      ↩︎ Start from the number of embeddings
      “One vector search unit covers up to 64 million vectors of dimension 768 (or the equivalent).”
      ↩︎ Start from the number of embeddings
      “If you anticipate low QPS across all indices, you can combine your indices under a single endpoint to avoid multiple base endpoint costs.”
      ↩︎ How endpoints are billed, and how to pay less
      “one endpoint can serve up to 50 indices”
      ↩︎ How endpoints are billed, and how to pay less
      “One vector search unit covers up to 2 million vectors of dimension 768 (or the equivalent).”
      ↩︎ Key concept
      “Endpoints with indexes that receive no query traffic still incur serving costs.”
      ↩︎ Exam trap 1
      “if you have 1 million vectors of dimension 1536, that also counts as one unit.”
      ↩︎ Prediction
      “AI Search endpoints are charged only after an index has been created”
      ↩︎ Checkpoint
    2. 2.
      “Choose standard endpoints when latency is critical and your index is well under 320M vectors.”
      ↩︎ Trade latency against cost per vector
      “need better cost efficiency per vector (up to 7x cheaper)”
      ↩︎ Trade latency against cost per vector
      “Eventually, QPS plateaus at approximately 30 QPS (ANN).”
      ↩︎ Trade latency against cost per vector
      “That makes 384 dimensions and below especially attractive for high-throughput, latency-sensitive use cases”
      ↩︎ Dimension is a sizing lever too
      “Conversely, higher-dimensional vectors increase compute load and reduce throughput.”
      ↩︎ Exam trap 2
      “Performance is highest when your index fits within a single vector search unit, with extra space to handle additional query load.”
      ↩︎ Checkpoint
      “typically improves QPS by about 1.5x and reduces latency by about 20%”
      ↩︎ Prediction
      “As a general rule, choose the smallest dimensionality that preserves retrieval quality for your use case.”
      ↩︎ Checkpoint
    3. 3.
      “For storage-optimized endpoints, the embedding dimension must be evenly divisible by 16.”
      ↩︎ Dimension is a sizing lever too

    Also cited

    Continue to page 2 of 2

    Vector Search Sync Modes, High QPS and Query Latency Tuning

    Spotted a mistake, or was something unclear? Tell us.