What you will be able to do
- Estimate how many vector search units an index needs from its vector count and dimension
- Choose between standard and storage-optimized endpoints from vector count, latency target and cost per vector
- Explain how embedding dimensionality affects latency and QPS
- Describe how AI Search endpoints are billed and how to keep the base cost down
Key concept
Vector search unit (VSU) — The unit of capacity that sizes and bills an AI Search endpoint. A standard VSU holds about 2M vectors at dimension 768, and a storage-optimized VSU holds about 64M. An endpoint grows by whole units as its indexes grow, so vector count and dimension together decide capacity, performance and cost.
1.Start from the number of embeddings
Databricks AI Search was formerly called Databricks Vector Search, and the exam may use either name. When you configure it for a solution, the first thing to settle is the endpoint type, and you pick it when you create the endpoint. There are two: Standard and Storage Optimized. The main input to that choice is how many embeddings you need to store.
Each type is measured in vector search units (VSUs). One standard VSU holds up to 2 million vectors at dimension 768. One storage-optimized VSU holds up to 64 million vectors at dimension 768. The phrase "or the equivalent" matters: what counts is total vector data, not just the row count.
Endpoint capacity sets a hard limit on the choice. A standard endpoint holds 320 million vectors at dimension 768. A storage-optimized endpoint holds over one billion and indexes 10–20x faster. Both endpoint types have a base price and scale up automatically to match the total size of the indexes they serve. They scale back down when an index is deleted, and the smallest possible endpoint is one VSU.
You set the type at creation, so make the decision before you build anything:
client.create_endpoint(
name="vector_search_endpoint_name",
endpoint_type="STANDARD" # or "STORAGE_OPTIMIZED"
)Checkpoint 1 of 5· Check yourself
A team expects about 600 million embeddings of dimension 768. Which endpoint type can hold them?
A standard endpoint holds 320M vectors at dimension 768, so 600M exceeds it. Storage-optimized endpoints hold over a billion.
“Storage-optimized endpoints have a larger capacity (over one billion vectors at dimension 768) and provide 10-20x faster indexing.”Source: docs.databricks.com
Sources1
2.Trade latency against cost per vector
Capacity is only part of the decision. The two types also have very different latency. The performance guide states the rule directly: choose standard when latency is critical and the index is well under 320M vectors. Choose storage-optimized when you have 10M+ vectors, can accept extra latency, and need better cost efficiency per vector, which can be up to 7x cheaper. The reference numbers below were measured with self-managed embeddings at dimension 768.
| SKU | Vectors | Latency | QPS |
|---|---|---|---|
| Standard | 10K | 20ms | 200+ |
| Standard | 10M | 40ms | 30 |
| Standard | 100M | 50ms | 30 |
| Storage-optimized | 10M | 300ms | 50 |
| Storage-optimized | 100M | 400ms | 40 |
| Storage-optimized | 1B | 500ms | 30 |
Notice what happens as standard indexes grow. At 10K vectors, QPS is 200+, but by 10M it has fallen to 30. Performance is best when the index fits inside one VSU with room to spare. Once it grows past one unit (2M+ vectors on standard, 64M+ on storage-optimized), latency rises and QPS levels off at around 30 for ANN queries. So vector count shapes performance even within one endpoint type.
Checkpoint 2 of 5· Match them up
Match each endpoint characteristic to its value
Tap a term, then the definition that fits it.
Standard endpoints have low latency and small units. Storage-optimized endpoints have large units and higher latency, and in return cost less per vector.
“Performance is highest when your index fits within a single vector search unit, with extra space to handle additional query load.”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
A team is building a RAG chatbot over a product catalog that changes every few seconds. Business stakeholders require the index to reflect catalog updates within seconds and are willing to accept higher compute cost for this freshness. Which vector search configuration best meets this requirement?
Correct answer: A — A Standard endpoint with Continuous sync mode, since it keeps the index updated with seconds of latency by running a provisioned streaming pipeline continuously.
- A. Continuous sync mode is only available on Standard endpoints and runs a provisioned streaming pipeline that keeps the index in sync with seconds of latency, which matches the near-real-time requirement even though it costs more than triggered syncing.
- B. Storage Optimized endpoints only support Triggered sync, which is user- or schedule-initiated and cannot deliver the seconds-level freshness the stakeholders require between manual sync calls.
- C. Hourly triggered syncs only reprocess changed rows since the last sync, but the hour-long gap between syncs leaves the index stale far longer than the required seconds-level freshness.
- D. A Direct Vector Access index requires the application itself to issue manual write calls for every change, which places the burden of tracking every catalog update on custom code rather than providing managed seconds-level freshness.
Sources2
3.Dimension is a sizing lever too
Embedding dimensions are usually 384, 768, 1024 or 1536. Higher dimensions can capture more meaning and improve retrieval quality, but each query costs more compute. Dimension also affects the count side of the decision, because a VSU is defined at dimension 768 or its equivalent. Larger vectors fill units faster.
The guide's rule is to use the smallest dimension that keeps retrieval quality acceptable. That makes 384 dimensions and below especially attractive for high-throughput, latency-sensitive applications. One constraint is specific to storage-optimized endpoints: if you supply precomputed embeddings, the dimension must be evenly divisible by 16.
Checkpoint 4 of 5· Check yourself
A latency-sensitive RAG app gets the same retrieval quality from a 384-dimension model as from a 1536-dimension model. What does the performance guide recommend?
Lower dimensions need less computation, which gives faster queries and higher QPS. When quality is equal, the smaller dimension wins.
“As a general rule, choose the smallest dimensionality that preserves retrieval quality for your use case.”Source: docs.databricks.com
4.How endpoints are billed, and how to pay less
AI Search bills for two things: the indexes that store your vectors and the endpoints that serve queries. An endpoint has a base price and then grows in whole VSUs. Billing starts only once an index exists on the endpoint. After the last index is deleted, charges stop 24 hours later.
Because every endpoint carries its own base price, consolidation is the main cost lever. One endpoint can serve up to 50 indexes. If you expect low QPS across all of them, putting them on one endpoint means you pay one base price instead of several. The reverse also applies: endpoints whose indexes get no query traffic still incur serving costs, so find them and remove them. You can track spending with the system.billing.usage table (filtered to billing_origin_product = 'VECTOR_SEARCH'), the usage dashboards, and usage policies that tag endpoint and index costs by team or project.
Move all four indexes onto a single endpoint. One endpoint serves up to 50 indexes, so you pay one base endpoint cost instead of four. Usage policies can still tag each team's costs separately.
Checkpoint 5 of 5· Check yourself
An engineer creates a standard endpoint on Monday and adds the first index on Thursday. When does the endpoint start being charged?
Endpoints are charged only after an index exists. Charges stop 24 hours after the last index is deleted.
“AI Search endpoints are charged only after an index has been created”Source: docs.databricks.com
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.An AI Search endpoint starts billing as soon as it is created, and an idle endpoint costs nothing.Why is that wrong?
Billing starts only after an index is created on the endpoint. From then on, an endpoint whose indexes get no queries still incurs serving costs until the indexes are removed.
2.Higher-dimensional embeddings are always the better configuration.Why is that wrong?
Higher dimensions add compute per query and fill VSUs faster. Use the smallest dimension that keeps retrieval quality acceptable.
Covered in Dimension is a sizing lever too
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“each endpoint has a base price and scales up automatically to match the total size of the indices it is serving.”
↩︎ Start from the number of embeddings“One vector search unit covers up to 64 million vectors of dimension 768 (or the equivalent).”
↩︎ Start from the number of embeddings“If you anticipate low QPS across all indices, you can combine your indices under a single endpoint to avoid multiple base endpoint costs.”
↩︎ How endpoints are billed, and how to pay less“one endpoint can serve up to 50 indices”
↩︎ How endpoints are billed, and how to pay less“One vector search unit covers up to 2 million vectors of dimension 768 (or the equivalent).”
↩︎ Key concept“Endpoints with indexes that receive no query traffic still incur serving costs.”
↩︎ Exam trap 1“if you have 1 million vectors of dimension 1536, that also counts as one unit.”
↩︎ Prediction“AI Search endpoints are charged only after an index has been created”
↩︎ Checkpoint - 2.
“Choose standard endpoints when latency is critical and your index is well under 320M vectors.”
↩︎ Trade latency against cost per vector“need better cost efficiency per vector (up to 7x cheaper)”
↩︎ Trade latency against cost per vector“Eventually, QPS plateaus at approximately 30 QPS (ANN).”
↩︎ Trade latency against cost per vector“That makes 384 dimensions and below especially attractive for high-throughput, latency-sensitive use cases”
↩︎ Dimension is a sizing lever too“Conversely, higher-dimensional vectors increase compute load and reduce throughput.”
↩︎ Exam trap 2“Performance is highest when your index fits within a single vector search unit, with extra space to handle additional query load.”
↩︎ Checkpoint“typically improves QPS by about 1.5x and reduces latency by about 20%”
↩︎ Prediction“As a general rule, choose the smallest dimensionality that preserves retrieval quality for your use case.”
↩︎ Checkpoint - 3.
“For storage-optimized endpoints, the embedding dimension must be evenly divisible by 16.”
↩︎ Dimension is a sizing lever too
Also cited
“Storage-optimized endpoints have a larger capacity (over one billion vectors at dimension 768) and provide 10-20x faster indexing.”
↩︎ Checkpoint