What you will be able to do
- Decide when reserved provisioned throughput is a better choice than pay-per-token, and what happens to overflow traffic
- Find out which models, services and users drive foundation model spend, using system.billing.usage and system.ai_gateway.external_model_spend
- Choose service tags or request tags so spend can be attributed to a team or project
1.Pay-per-token or reserved provisioned throughput
How you buy model capacity is a cost decision in its own right. Foundation Model APIs default to pay-per-token, where you pay for each request. Reserved provisioned throughput works the other way round: you prepay for a fixed pool of model units on one foundation model for a set term, and Databricks backs that pool with dedicated capacity for the length of the reservation. More model units means more guaranteed throughput. Databricks provides a model units estimator to help you size the pool from the traffic you expect. One endpoint can hold several reservations at the same time, and each has its own model units and term-end date.
Databricks recommends reserved provisioned throughput for business-critical applications or agents that need guaranteed throughput and latency and have predictable traffic. If traffic is unpredictable, the prepaid pool may sit idle, or much of the load may spill over to pay-per-token anyway. Not every model can be reserved: Databricks decides eligibility per model, and the create flow shows only the models you can reserve. Endpoint settings also matter. Scale to zero cuts resource use to zero when an endpoint is idle, and Databricks recommends it for testing and development. It is not recommended for production, because latency is higher and capacity is not guaranteed while the endpoint is scaled to zero.
Checkpoint 1 of 4· Check yourself
Which workload best fits reserved provisioned throughput?
Reservations make sense when traffic is predictable and you need reliable throughput and latency. Development endpoints that are idle most of the time are a better fit for scale to zero.
“guaranteed, reliable throughput and latency, and your traffic is predictable”Source: docs.databricks.com
2.Finding out what drives spend: billing and spend tables
Before you can control costs, you need to know where they come from. Unity Gateway adds extra fields to MODEL_SERVING records in system.billing.usage so that cost can be traced back. usage_metadata.ai_gateway.endpoint_name holds the model service name, usage_metadata.ai_gateway.destination_model holds the model that served the request, identity_metadata.run_by holds the user or service principal that sent it, and custom_tags holds the service tags. This applies to both real-time and batch requests that go through the gateway. These records are in DBUs. To turn them into dollars, join system.billing.list_prices. Those dollar figures are at list price, so they don't include negotiated discounts or credits.
SELECT
usage_metadata.ai_gateway.destination_model AS destination_model,
SUM(usage_quantity) AS dbus
FROM system.billing.usage
WHERE billing_origin_product = 'MODEL_SERVING'
AND usage_metadata.ai_gateway.endpoint_name IS NOT NULL
AND usage_unit = 'DBU'
AND usage_date >= current_date() - INTERVAL 30 DAYS
GROUP BY destination_model
ORDER BY dbus DESC;External models are tracked separately. For requests routed through model provider services, Databricks estimates the dollar cost from token usage and the provider's published prices, applies any price multiplier configured on the service, and writes hourly totals to system.ai_gateway.external_model_spend. In that table, usage_unit is always USD, so you can sum usage_quantity directly with no price join. These figures are for information only and may not match the provider's invoice. If you'd rather not write SQL, the built-in usage dashboard (version 0.4 and above) has a Cost Analysis page covering both Databricks and external spend.
Checkpoint 2 of 4· Fill the gap
Which table completes this query, which ranks users by estimated external-provider spend in USD?
SELECT
identity_metadata.run_by AS run_by,
SUM(usage_quantity) AS usd
FROM ?
WHERE usage_start_time >= current_timestamp() - INTERVAL 30 DAYS
GROUP BY run_by
ORDER BY usd DESC;External model spend is stored as hourly USD estimates in system.ai_gateway.external_model_spend. system.billing.usage covers Databricks-hosted models and records DBUs.
Source: docs.databricks.com3.Attributing spend to teams and projects with tags
run_by tells you which user spent the money, but not which team or project it was for. Tags fill that gap, and there are two kinds. Service (endpoint) tags go on a model service, for example team = ml-platform, and every request through that service inherits them. Use these when each team or project has its own service. Request tags are set by the caller on each request with an HTTP header. Use these when many teams share one service.
Databricks-Ai-Gateway-Request-Tags: {"project": "chatbot", "team": "ml-platform"}| Tag type | system.ai_gateway.usage | system.billing.usage | system.ai_gateway.external_model_spend |
|---|---|---|---|
| Service tags | endpoint_tags | custom_tags | Yes |
| Request tags | request_tags | Not propagated | Yes |
No. For Databricks-provided models, only service tags reach system.billing.usage. Request tags still appear in system.ai_gateway.usage, so you can split request counts and token totals by team there. To split billed cost, give each team its own model service with a service tag.
SELECT
endpoint_tags['team'] AS team,
endpoint_name,
destination_model,
COUNT(*) AS request_count,
SUM(total_tokens) AS total_tokens
FROM system.ai_gateway.usage
WHERE endpoint_tags['team'] IS NOT NULL
GROUP BY endpoint_tags['team'], endpoint_name, destination_model
ORDER BY total_tokens DESC;Checkpoint 3 of 4· Check yourself
You need billed cost for an external provider model, broken down by project, on a service shared by many projects. Which approach works?
For external provider models, both service tags and request tags reach system.ai_gateway.external_model_spend, so per-request tags can split a shared service.
“External provider models: use service or request tags. Both propagate to system.ai_gateway.external_model_spend, which captures external model spend estimates.”Source: docs.databricks.com
Checkpoint 4 of 4· Exam question
A platform team supports five separate agent applications, each tagged with its own cost-center identifier, all routed through a shared Unity Catalog AI Gateway endpoint. Finance wants an email alert the moment any one cost center's monthly spend on that endpoint crosses $2,000, and they want the option to automatically block further requests from that cost center once the threshold is hit. Which Databricks feature should the team configure?
Correct answer: A — A budget scoped to the cost-center tags with a defined threshold and usage blocking enabled for the AI Gateway endpoint.
- A. Databricks budgets can be scoped to custom tags such as cost-center identifiers, tracking spend against a defined monthly threshold and sending near real-time alerts for Unity AI Gateway usage, with an optional usage-blocking setting that halts further requests once the threshold is reached, matching both the alerting and blocking requirements.
- B. A tokens-per-minute rate limit caps request throughput rather than tracking dollar spend against a threshold, and applying one limit uniformly across all cost centers would not distinguish which specific cost center crossed $2,000, so it does not satisfy the tag-based spend alerting requirement.
- C. An inference table logs request and response payloads for auditing, evaluation, and monitoring quality, but it has no built-in mechanism to compute dollar spend per tag, trigger a threshold alert, or block further calls.
- D. An access control list can permit or deny which principals may query an endpoint, but it does not track cumulative spend or generate a threshold-crossing alert; it would only enforce an all-or-nothing access decision, not a cost-based one.
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Request tags are enough to attribute DBU cost for Databricks-provided models in system.billing.usage.Why is that wrong?
Only service tags propagate to system.billing.usage. Request tags appear in system.ai_gateway.usage and in external model spend data.
Covered in Attributing spend to teams and projects with tags
2.Reserved provisioned throughput caps your bill at the cost of the reservation.Why is that wrong?
Traffic above the reserved pool spills over to priority pay-per-token and is billed as that.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/reserved-provisioned-throughputOfficial docs
“you reserve a fixed pool of model units on a single foundation model for a set term”
↩︎ Pay-per-token or reserved provisioned throughput“Databricks backs that pool with dedicated capacity for the length of the reservation.”
↩︎ Pay-per-token or reserved provisioned throughput“More model units means more guaranteed throughput.”
↩︎ Pay-per-token or reserved provisioned throughput“If your traffic exceeds the reserved pool, the overflow is served automatically on priority pay-per-token and billed accordingly.”
↩︎ Exam trap 2“If your traffic exceeds the reserved pool, the overflow is served automatically on priority pay-per-token and billed accordingly.”
↩︎ Prediction“guaranteed, reliable throughput and latency, and your traffic is predictable”
↩︎ Checkpoint - 2.
“scale to zero is not recommended for production endpoints, as latency is greater and capacity is not guaranteed when scaled to zero.”
↩︎ Pay-per-token or reserved provisioned throughput - 3.
“Unity Gateway provides cost attribution through the billable usage system table (system.billing.usage).”
↩︎ Finding out what drives spend: billing and spend tables“Spend is aggregated hourly and recorded in the system.ai_gateway.external_model_spend system table.”
↩︎ Finding out what drives spend: billing and spend tables - 4.
“They exclude negotiated discounts and billing credits.”
↩︎ Finding out what drives spend: billing and spend tables“Request tags: when many teams or projects share the same model service.”
↩︎ Attributing spend to teams and projects with tags“Databricks-provided models: use service tags. Only service tags propagate to system.billing.usage.”
↩︎ Attributing spend to teams and projects with tags“Databricks-provided models: use service tags. Only service tags propagate to system.billing.usage.”
↩︎ Exam trap 1“External provider models: use service or request tags. Both propagate to system.ai_gateway.external_model_spend, which captures external model spend estimates.”
↩︎ Checkpoint