What you will be able to do
- Identify which AI Gateway features apply to an agent deployed via Agent Framework, and which of them are billed
- Enable inference tables for an agent endpoint and query them, including joining to system.serving.served_entities
- Explain what the system.serving usage tables record, who can read them, and how token counts are estimated
- Configure endpoint, default-user, custom-user and group rate limits, and predict which limit is enforced
Key concept
AI Gateway on a serving endpoint — AI Gateway is a governance layer on a Model Serving endpoint. You switch on its features one at a time: payload logging to inference tables, usage tracking in system tables, and rate limits. Each feature's data ends up in Unity Catalog, where you can query it to see what your deployed LLM or agent is doing.
1.What AI Gateway can track on an agent endpoint
When you deploy an agent with Agent Framework, it runs behind a Model Serving endpoint. AI Gateway is configured on that endpoint, not inside the agent code. On the endpoint creation page, the AI Gateway section lets you turn each feature on separately. The features this lesson covers are payload logging (inference tables), usage tracking (system tables), and rate limiting. All of their data is logged into Delta tables in Unity Catalog, so you analyse it with ordinary SQL or notebooks.
| Feature | What it does | Databricks agents | Foundation Model APIs pay-per-token endpoint |
|---|---|---|---|
| Payload logging | Monitor and audit data being sent to model APIs using inference tables | Supported | Supported |
| Usage tracking | Monitor operational usage on endpoints and associated costs using system tables | Not supported | Supported |
| Permission and rate limiting | Control who has access and how much access | Not supported | Supported |
| AI Guardrails | Prevent unwanted and unsafe data in requests and responses | Not supported | Supported |
| Available in Unity Gateway | Use enhanced Unity Gateway features | Not supported | Supported |
Read the table carefully. Inference tables are the AI Gateway feature you can count on for an agent endpoint. The matrix marks usage tracking and rate limiting as Not supported for Databricks agents. However, the rate-limit configuration guide only says that *TPM* limits cannot be applied to agent endpoints (see the rate-limiting section below). The sources don't reconcile these two statements, so treat 'inference tables for agents' as the firm fact. The usage tables and rate limits are fully supported on the Foundation Model APIs, external model and custom model endpoints. The newer Unity Gateway experience is also marked Not supported for agents, so this lesson uses the serving-endpoint version of AI Gateway throughout.
Cost matters too. AI Gateway charges per feature you enable. The two tracking features are billed, and the controls are free.
Checkpoint 1 of 8· Check yourself
A team enables inference tables, usage tracking and rate limits on a serving endpoint. Which of these add AI Gateway charges?
The docs name inference tables and usage tracking as the paid features. Rate limiting, query permissions, fallbacks and traffic splitting are free.
“Paid features include inference tables and usage tracking.”Source: docs.databricks.com
2.Turning on inference tables for an agent
An inference table captures the requests coming into an endpoint and the responses going out, and logs them as a Unity Catalog Delta table. Agent endpoints get one extra thing: the agent's inference table stores payload and request details as well as MLflow Trace logs. That means you can follow each step the agent took, not only its final answer.
How you enable it depends on how you deployed:
- **Agents deployed with mlflow.deploy() have inference tables enabled automatically.
- Programmatic deployments** need the ENABLE_MLFLOW_TRACING environment variable set to True in the endpoint configuration.
- Any endpoint in the Serving UI can have them switched on in the AI Gateway section, either when you create it or later with *Edit AI Gateway*. Clearing the *Enable inference table* checkbox there disables them.
Checkpoint 2 of 8· Put it in order
Put the UI steps for enabling inference tables while creating a new endpoint in order
- 1.In the AI Gateway section, select Enable inference tables
- 2.Click Create serving endpoint
- 3.Click Serving in the Databricks UI
Start in Serving, create the endpoint, then tick the checkbox in that endpoint's AI Gateway section.
“In the AI Gateway section, select Enable inference tables.”Source: docs.databricks.com
Some prerequisites catch people out. The workspace needs Unity Catalog and serverless compute enabled. Both the person who creates the endpoint and anyone who modifies it need Can Manage on the endpoint. In Unity Catalog they also need USE CATALOG, USE SCHEMA and CREATE TABLE on the target catalog and schema. You can't use an OpenSharing catalog as the target.
Databricks creates the table for you, and you can't point it at an existing table. After that, leave the table alone. If you change its schema, rename it or delete it, logging can stop or the table can become corrupted. The user who enabled the inference table owns it, and access is governed by standard Unity Catalog ACLs.
Logging is not real-time. Rows appear less than an hour after the endpoint is queried. Payloads larger than 1 MiB are not logged. Streaming responses are supported: the logged response combines all the chunks returned.
3.Reading an inference table to monitor the agent
Once the endpoint is serving, every request and its response go into the inference table automatically. You can open the table from the endpoint page, which takes you to Catalog Explorer. You can also query it from Databricks SQL or a notebook, or through the REST API. The docs list four common uses:
1. Debug production issues. The table records HTTP status codes, request and response JSON, model run times and traces. 2. Monitor data and model quality with data profiling dashboards and alerts. 3. Build a training corpus by joining the logs with ground truth labels. 4. Monitor deployed agents through the MLflow traces stored alongside the payloads.
| Column | Type | What it tells you |
|---|---|---|
| databricks_request_id | STRING | Databricks-generated identifier attached to every model serving request |
| client_request_id | STRING | Your own request identifier, if you passed one in the request body |
| status_code | INT | HTTP status code returned from the model |
| execution_duration_ms | BIGINT | Time the model spent on inference, excluding network overhead |
| request / response | STRING | Raw JSON bodies sent to and returned by the endpoint |
| sampling_fraction | DOUBLE | Between 0 and 1; 1 means 100% of requests were included |
| logging_error_codes | ARRAY | Why a row couldn't be logged, e.g. MAX_REQUEST_SIZE_EXCEEDED |
| requester | STRING | User or service principal whose permissions were used for the call |
execution_duration_ms counts only the time the model spent generating predictions. It leaves out network overhead, so latency that users experience outside the model doesn't appear in this column.
The inference table holds a served_entity_id but nothing about what that entity actually is. Those details live in the system.serving.served_entities system table, so you join the two on served_entity_id.
Checkpoint 3 of 8· Fill the gap
Which system table completes this join that adds details about the served model to each logged request?
SELECT * FROM <catalog>.<schema>.<payload_table> payload
JOIN system.serving. ? se on payload.served_entity_id = se.served_entity_idFoundation model details are stored in system.serving.served_entities, which is keyed by served_entity_id. endpoint_usage holds per-request token counts, not entity metadata.
Source: docs.databricks.comSources3
4.Usage tables: token counts and endpoint metadata
Inference tables answer the question *what was said*. Usage tracking answers *how much was used, and by whom*. To turn it on, select Enable usage tracking in the AI Gateway section. Unity Catalog must be enabled. Two system tables are then shared automatically:
| System table | Records | Example columns |
|---|---|---|
| system.serving.endpoint_usage | Token counts for each request to the endpoint | requester, status_code, request_time, input_token_count, output_token_count, usage_context |
| system.serving.served_entities | Metadata for each served entity | served_entity_id, endpoint_name, entity_type, entity_name, entity_version, task |
A few details are commonly tested:
- Token counts can be estimates. If the model doesn't return a token count, AI Gateway estimates input and output tokens as (text_length+1)/4.
- Custom models show zero. input_token_count, output_token_count and the character-count columns are 0 for custom model requests.
- Attribution. usage_context is a user-provided map that identifies the end user or customer application making the call. Use it to attribute cost beyond the requester identity.
- Entity types. In served_entities, entity_type can be FEATURE_SPEC, EXTERNAL_MODEL, FOUNDATION_MODEL or CUSTOM_MODEL. For pay-per-token endpoints, created_by is System-User.
Remember from the support matrix that usage tracking is marked Not supported for Databricks agents, and Supported for external model, Foundation Model APIs and custom model endpoints. The sources don't explain how an agent's token consumption is attributed in these tables. The documented use is per-endpoint usage on the supported endpoint types.
Checkpoint 4 of 8· Check yourself
An ML engineer with CAN MANAGE on an endpoint enables usage tracking. Who can query system.serving.endpoint_usage by default?
The endpoint manager must enable usage tracking, but only account admins can view or query these tables. Admins can grant access to system tables to other users.
“Only account admins have permission to view or query the served_entities table or endpoint_usage table”Source: docs.databricks.com
Checkpoint 5 of 8· Exam question
An admin configures a 100 QPM service-level rate limit on an AI Gateway-enabled endpoint, then adds a 20 QPM group limit for a data science group, and finally sets a 5 QPM custom limit for one user in that group who has been issuing an unusually high volume of test calls. When that specific user sends requests, which limit governs their traffic, and what happens once it is exceeded?
Correct answer: D — The user's 5 QPM custom limit takes priority over the group and service limits, and requests beyond it receive an HTTP 429 Too Many Requests response.
- A. Service-level limits act as an overall ceiling, but a more specific user-level limit still governs that individual's traffic, and AI Gateway rejects excess requests rather than queuing them.
- B. AI Gateway evaluates custom user-level limits ahead of group-level limits, not the reverse, and exceeded requests return HTTP 429 rather than HTTP 503.
- C. Precedence among configured rate limits is based on specificity (individual user over group over service), not recency of configuration, and the gateway does not auto-retry rejected calls.
- D. Correct: a custom limit scoped to an individual user takes priority over broader group and service-level limits, and once that user's requests exceed 5 QPM the gateway returns HTTP 429.
Sources2
5.Rate limits: capping how hard the endpoint can be hit
Rate limits come in two kinds: queries per minute (QPM) and tokens per minute (TPM). As the predict showed, TPM is off the table for agent and custom model endpoints. (The support matrix goes further and marks permission and rate limiting as Not supported for Databricks agents. Check which endpoint type a question describes.) Rate limits only apply to users who have permission to query the endpoint, and by default none are configured. You can set limits at four levels:
| Level | Scope | Precedence |
|---|---|---|
| Endpoint | All traffic to the endpoint, regardless of user | Global maximum: if exceeded, all requests are blocked regardless of user or group limits |
| User (Default) | Every user of the endpoint without a more specific limit | Overridden by any custom rate limit |
| Custom: individual user or service principal | One identity | Takes priority over user group custom limits |
| Custom: user group | Shared limit for all members of the group | User-specific limit wins when both apply |
Three more rules complete the picture:
- When an endpoint, user or service principal has both a QPM and a TPM limit, the more restrictive one is enforced. - When a user belongs to several groups with different limits, they are throttled only once they exceed all of their groups' QPM limits, or all of their groups' TPM limits. - An endpoint can have at most 20 rate limits, of which at most 5 are group-specific.
Rate limits are only as strong as your endpoint permissions. To stop people bypassing throughput limits, restrict endpoint creation and CAN MANAGE to admins, and give everyone else query permission only on approved endpoints.
Checkpoint 6 of 8· Match them up
Match each scenario to the limit that governs it
Tap a term, then the definition that fits it.
The endpoint limit is a global ceiling. A user-specific limit beats a group limit, any custom limit overrides the default, and when QPM and TPM both apply, the stricter one wins.
“If a user belongs to both a user-specific limit and a group-specific limit, the user-specific limit is enforced.”Source: docs.databricks.com
Checkpoint 7 of 8· Check yourself
Bob is in two groups: Analysts (100 QPM) and Interns (20 QPM). He has no user-specific limit. When is Bob rate limited on QPM?
With several groups, the user is throttled only after exceeding every group's QPM limit, so effectively at the most generous one. The 'more restrictive wins' rule applies to QPM versus TPM, not to multiple groups.
“the user is rate limited if they exceed all of the QPM rate limits or all of the TPM rate limits”Source: docs.databricks.com
Checkpoint 8 of 8· Exam question
An engineer configures both a 50 QPM limit and a 5,000 TPM limit on the same pay-per-token endpoint serving an agent. During a traffic burst, most incoming requests use very few tokens each, so the request count crosses 50 QPM well before the token budget is anywhere near 5,000 TPM. What does AI Gateway do in this situation?
Correct answer: A — It enforces whichever configured limit is reached first, so it returns HTTP 429 once the 50 QPM threshold is crossed even though the token budget is unused.
- A. Correct: when both request-based and token-based limits are configured, AI Gateway enforces whichever limit is more restrictive in practice, so hitting the QPM ceiling first triggers HTTP 429 regardless of remaining token budget.
- B. AI Gateway does not require both thresholds to be breached simultaneously; crossing either configured limit on its own is enough to trigger rejection of further requests.
- C. Token-based limits do not automatically take precedence over request-based limits; both are evaluated, and the more restrictive one in effect at that moment is what gets enforced.
- D. AI Gateway does not blend QPM and TPM into a single combined threshold; the two limits are tracked and enforced independently, with the more restrictive one taking effect.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.You can cap an agent endpoint's cost with a tokens-per-minute (TPM) rate limit.Why is that wrong?
TPM limits can't be applied to endpoints that serve custom models or agents. Token-based limits belong to LLM endpoints such as external models and Foundation Model APIs.
Covered in Rate limits: capping how hard the endpoint can be hit
2.A group limit overrides a user's personal limit, because groups are managed centrally.Why is that wrong?
When a user has both a user-specific limit and a group limit, the user-specific limit is enforced.
Covered in Rate limits: capping how hard the endpoint can be hit
3.Token counts in the usage table always come straight from the model, so they are exact.Why is that wrong?
If the model doesn't return a token count, AI Gateway estimates input and output tokens from the text length.
4.You can point an agent's inference logging at an existing Delta table you've already prepared.Why is that wrong?
Databricks always creates a new inference table itself. Changing that table's schema or name, or deleting it, can stop logging or corrupt it.
Covered in Turning on inference tables for an agent
5.Rate limiting is a billed AI Gateway feature, like inference tables.Why is that wrong?
Only the tracking features (inference tables and usage tracking) are billed. Rate limiting, query permissions, fallbacks and traffic splitting are free.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“All data is logged into Delta tables in Unity Catalog.”
↩︎ What AI Gateway can track on an agent endpoint“AI Gateway incurs charges on an enabled feature basis.”
↩︎ What AI Gateway can track on an agent endpoint“It is a centralized service that brings governance, monitoring, and production readiness to model serving endpoints.”
↩︎ Key concept“Features such as query permissions, rate limiting, fallbacks, and traffic splitting are free of charge.”
↩︎ Exam trap 5“Paid features include inference tables and usage tracking.”
↩︎ Checkpoint - 2.
“In the AI Gateway section of the endpoint creation page, you can individually configure Unity Gateway features.”
↩︎ What AI Gateway can track on an agent endpoint“Payload logging data populates these tables less than an hour after querying the endpoint.”
↩︎ Turning on inference tables for an agent“Payloads larger than 1 MiB are not logged.”
↩︎ Turning on inference tables for an agent“system.serving.endpoint_usage, which captures token counts for each request to the endpoint.”
↩︎ Usage tables: token counts and endpoint metadata“The token count of the input. This will be 0 for custom model requests.”
↩︎ Usage tables: token counts and endpoint metadata“The user provided map containing identifiers of the end user or the customer application”
↩︎ Usage tables: token counts and endpoint metadata“Rate limits only apply to users who have permission to query the endpoint.”
↩︎ Rate limits: capping how hard the endpoint can be hit“The endpoint rate limit is a global maximum.”
↩︎ Rate limits: capping how hard the endpoint can be hit“A maximum of 20 rate limits and up to 5 group-specific rate limits can be specified on an endpoint.”
↩︎ Rate limits: capping how hard the endpoint can be hit“the more restrictive rate limit is enforced”
↩︎ Rate limits: capping how hard the endpoint can be hit“To prevent bypassing guardrails or throughput limits, restrict endpoint creation and CAN MANAGE to admins”
↩︎ Rate limits: capping how hard the endpoint can be hit“TPM rate limits cannot be applied to serving endpoints that serve custom models or agents.”
↩︎ Exam trap 1“If a user belongs to both a user-specific limit and a group-specific limit, the user-specific limit is enforced.”
↩︎ Exam trap 2“The input and output token count are estimated as (text_length+1)/4 if the token count is not returned by the model.”
↩︎ Exam trap 3“Only account admins have permission to view or query the served_entities table or endpoint_usage table”
↩︎ Checkpoint“TPM rate limits cannot be applied to serving endpoints that serve custom models or agents.”
↩︎ Prediction“If a user belongs to both a user-specific limit and a group-specific limit, the user-specific limit is enforced.”
↩︎ Checkpoint“the user is rate limited if they exceed all of the QPM rate limits or all of the TPM rate limits”
↩︎ Checkpoint - 3.
“these inference tables store payload and request details as well as MLflow Trace logs.”
↩︎ Turning on inference tables for an agent“Agents deployed using the mlflow.deploy() API have inference tables automatically enabled.”
↩︎ Turning on inference tables for an agent“For programmatic deployments, set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.”
↩︎ Turning on inference tables for an agent“Serverless compute has to be enabled in the workspace.”
↩︎ Turning on inference tables for an agent“Inference tables log data like HTTP status codes, request and response JSON code, model run times, and traces output during model run times.”
↩︎ Reading an inference table to monitor the agent“This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”
↩︎ Reading an inference table to monitor the agent“Foundation model details are captured in the system.serving.served_entities system table.”
↩︎ Reading an inference table to monitor the agent“Specifying an existing table is not supported.”
↩︎ Exam trap 4“Monitor deployed agents. Inference tables can also store MLflow traces for agents”
↩︎ Prediction“In the AI Gateway section, select Enable inference tables.”
↩︎ Checkpoint