What you will be able to do
- Identify which LLM spend a Unity Gateway budget covers and which it does not, including provisioned throughput
- Configure shared, per-user and override thresholds with alert and block actions
- Explain why budget enforcement is approximate and which table is the billing source of truth
- Choose service, per-user and custom rate limits (QPM/TPM) and predict what a client sees when one is exceeded
Key concept
Gateway-level spend guardrails — Databricks controls LLM cost at the Unity Gateway: budgets set a cap on monthly spend in dollars and can alert or block, and rate limits cap how many requests or tokens per minute a service accepts. Both apply only to traffic that flows through the gateway.
1.What a Unity Gateway budget actually covers
A budget is the most direct cost control Databricks offers for generative AI. It watches billing records and compares monthly spend against thresholds you set. You can scope a budget to the Unity Gateway product alone, or you can include Unity Gateway in a budget for the whole account. If the budget is scoped to Unity Gateway, enforcement is near real-time: once spend crosses a threshold, the alerts and blocking take effect within moments.
| Traffic type | Counted by the budget? |
|---|---|
| pay-per-token usage | Yes |
| ai_query | Yes |
| Genie | Yes |
| External model providers (model provider services) | Only if Include external model usage is selected |
| Provisioned throughput | No |
| Traffic sent directly to standalone Model Serving endpoints, bypassing Unity Gateway | No |
External models need a deliberate choice. When you create the budget you can select an option that will include estimated spend from requests routed to external models through model provider services. Once included, this spend counts toward the same shared and per-user thresholds. Databricks works out this estimate from token usage and the provider's published prices, so it can differ from the invoice the provider sends you. If you don't select the option, the budget still counts Databricks-hosted model usage.
Checkpoint 1 of 5· Check yourself
A workspace admin wants a budget to count spend from requests sent to an external provider through a model provider service. What must they do?
External model spend is opt-in per budget. If the option isn't selected, the budget counts only Databricks-hosted models.
“select Include external model usage to include estimated spend from requests routed to external models through model provider services.”Source: docs.databricks.com
Sources1
2.Thresholds and the two actions: alert or block
With the scope settled, you decide how the limit is split up. A shared threshold is one combined limit for every user in the budget scope. A per-user threshold gives each individual user their own monthly limit. A per-user override gives particular users or groups a higher limit than the default per-user threshold. If a user is in more than one override group, the highest of those thresholds is the one that applies. Each budget can have at most four shared thresholds and 20 per-user overrides. An account can have at most 1,000 budgets.
For each shared or per-user threshold you choose an action: Send alert, Block usage, or both. An alert sends email and the user keeps access. Block usage stops the user from sending any more requests through Unity Gateway and shows them a message saying their budget is exhausted. Access comes back when the budget resets or when an admin raises the threshold.
Checkpoint 2 of 5· Match them up
Match each budget setting to its effect
Tap a term, then the definition that fits it.
Thresholds decide whose spend is measured. Actions decide what happens when a threshold is crossed. Only Block usage takes access away.
“Send alert: Sends an email notification to configured addresses. Users retain access to Unity Gateway.”Source: docs.databricks.com
Sources1
3.Why a budget is not a hard spending cap
Enforcement works from a near-real-time estimate of cost, not from the final bill, so it is only approximate. A user may be blocked a little before they actually reach the threshold, or some spend may get through above it before the block takes effect. A request that is already running when the threshold is reached is allowed to finish. Databricks says plainly that budgets are a near-real-time cost control and not a guarantee.
The spend figures in different places can disagree, which confuses people. An alert email shows spend at the moment the alert fired. The budget details page updates often. system.billing.usage refreshes every few hours, and system.ai_gateway.external_model_spend refreshes hourly. If you query a system table just after an alert arrives, you can see a lower number. That delay doesn't affect enforcement, because blocking runs from its own near-real-time tracking.
Checkpoint 3 of 5· Check yourself
An alert email reports that a team crossed its threshold. Ten minutes later, system.billing.usage shows a lower total. What explains this?
Each source refreshes at a different rate. Enforcement doesn't depend on the system tables refreshing.
“Budget enforcement, including usage blocking and alert thresholds, is based on near real-time spend tracking and is not affected by system table refresh timing.”Source: docs.databricks.com
Sources1
4.Rate limits: capping throughput per service, user and group
A budget limits dollars over a month. A rate limit limits traffic per minute, and Databricks describes it as a way to manage both capacity and cost. Rate limits are part of a service's configuration. You can set them in the UI, or through the REST API, the SDKs, the CLI or Terraform. Model services accept requests-per-minute (QPM) and tokens-per-minute (TPM) limits. MCP services accept only QPM. No rate limits are configured by default.
| Level | Applies to | Precedence |
|---|---|---|
| Service | All traffic to the service, regardless of user | Global maximum; when exceeded, every request is blocked |
| User (Default) | Every user of the service without a custom limit | Overridden by custom rate limits |
| Custom: individual user or service principal | That principal | Takes priority over group custom limits |
| Custom: user group | Shared by all members of the group | Up to 5 group-specific limits per service |
Group membership adds a further rule. If a user has a user-specific limit and also belongs to a group with a limit, the user-specific limit wins. If a user belongs to several groups with different limits, they are limited only once they exceed all of their groups' QPM limits, or all of their groups' TPM limits. A rejected request gets HTTP 429 (Too Many Requests), and clients should retry with exponential backoff. The limiter is built for low latency, so it records usage after each response instead of checking concurrent requests in advance. Short bursts above the limit can happen, but over a longer window the average rate settles at the configured limit.
Checkpoint 4 of 5· Check yourself
An agent's client starts getting HTTP 429 responses from a Unity Gateway model service. What is the recommended client-side response?
A 429 means a rate limit was exceeded, not that a budget is exhausted. Databricks recommends retrying with exponential backoff.
“When a rate limit is exceeded, the service returns an HTTP 429 (Too Many Requests) response.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A generative AI engineer is serving a foundation model that receives a steady, predictable volume of production traffic 24/7, and the application's latency SLA requires consistently low and stable response times that cannot tolerate the variability of shared, on-demand capacity. Leadership also wants the ability to plan a fixed monthly cost for this workload rather than a variable bill tied to token volume. Which Databricks Foundation Model APIs deployment option best fits this scenario?
Correct answer: A — Provisioned throughput, sized with a minimum and maximum tokens-per-second range that matches the steady traffic pattern.
- A. Provisioned throughput allocates dedicated inference capacity sized in tokens per second, which delivers the consistent, predictable latency needed for steady 24/7 traffic and lets the team plan a fixed capacity cost rather than a per-token variable bill, matching both requirements in the scenario.
- B. Pay-per-token serving is well suited to variable or bursty workloads and bills per token consumed, so cost still fluctuates with volume; it does not offer the dedicated capacity needed to guarantee consistently stable latency under steady load.
- C. Scheduling `ai_query` as a periodic batch job is designed for offline or asynchronous bulk inference over tables of data, not for serving live requests with a low-latency SLA, so it does not fit an interactive production workload.
- D. Routing to a third-party provider's on-demand API through an external model endpoint still bills per request to that provider and does not give Databricks-managed dedicated capacity, so it does not provide the fixed-cost, guaranteed-latency behavior the scenario requires.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A Unity Gateway budget with Block usage will also stop spend on provisioned throughput endpoints.Why is that wrong?
Unity Gateway budgets don't track provisioned throughput. They also don't track traffic sent straight to standalone Model Serving endpoints without going through the gateway.
Covered in What a Unity Gateway budget actually covers
2.Block usage guarantees that billed spend never goes above the threshold.Why is that wrong?
Enforcement works from near-real-time estimates and lets in-flight requests finish, so final spend can exceed the threshold.
Covered in Why a budget is not a hard spending cap
3.A generous custom limit for a VIP user lets them exceed the service-level rate limit.Why is that wrong?
The service limit is a global maximum. Once it is exceeded, all requests are blocked, whatever user or group limits apply.
Covered in Rate limits: capping throughput per service, user and group
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“For budgets scoped to Unity Gateway, enforcement and alerts are near real-time.”
↩︎ What a Unity Gateway budget actually covers“Traffic sent directly to standalone Model Serving endpoints that doesn't flow through Unity Gateway”
↩︎ What a Unity Gateway budget actually covers“If a user belongs to multiple override groups, the highest threshold applies.”
↩︎ Thresholds and the two actions: alert or block“Each budget supports a maximum of four shared thresholds and 20 per-user overrides.”
↩︎ Thresholds and the two actions: alert or block“Access resumes after the budget resets or an admin increases the threshold.”
↩︎ Thresholds and the two actions: alert or block“Active requests in progress when the threshold is reached are not interrupted.”
↩︎ Why a budget is not a hard spending cap“system.billing.usage remains the source of truth for actual Databricks-billed usage.”
↩︎ Why a budget is not a hard spending cap“Budgets help you monitor and control monthly Unity Gateway spend using filters applied to billing records.”
↩︎ Key concept“Provisioned throughput inference is not currently tracked.”
↩︎ Exam trap 1“Do not use this feature as a way to ensure an absolute spend cap on final billed amounts.”
↩︎ Exam trap 2“Provisioned throughput inference is not currently tracked.”
↩︎ Prediction“select Include external model usage to include estimated spend from requests routed to external models through model provider services.”
↩︎ Checkpoint“Send alert: Sends an email notification to configured addresses. Users retain access to Unity Gateway.”
↩︎ Checkpoint“Budget enforcement, including usage blocking and alert thresholds, is based on near real-time spend tracking and is not affected by system table refresh timing.”
↩︎ Checkpoint - 2.
“Rate limits let you enforce consumption limits on a model service or MCP service to manage capacity and costs.”
↩︎ Rate limits: capping throughput per service, user and group“If a user belongs to both a user-specific limit and a group-specific limit, the user-specific limit is enforced.”
↩︎ Rate limits: capping throughput per service, user and group“Over a longer time window, the average request rate converges to the configured limit.”
↩︎ Rate limits: capping throughput per service, user and group“The service rate limit is a global maximum.”
↩︎ Exam trap 3“the more restrictive rate limit is enforced.”
↩︎ Prediction“When a rate limit is exceeded, the service returns an HTTP 429 (Too Many Requests) response.”
↩︎ Checkpoint