What you will be able to do
- Explain what an AI Gateway-enabled inference table captures and what it is used for
- Enable inference tables on a serving endpoint and name the permissions this requires
- Read the inference table schema to work out latency, status and requester for each request
- Spot the limits that leave gaps in an inference table: payload size, error codes, delivery delay and table changes
- Turn on inference tables with MLflow traces for a deployed agent
Key concept
Inference table — An inference table is a Unity Catalog Delta table. Databricks fills it automatically with every request a serving endpoint receives and every response it returns. Because it is ordinary governed data, you can query it with SQL to monitor, debug, audit and improve a live model or agent.
1.What an inference table gives you that health metrics do not
Once an LLM endpoint is live, there are two kinds of question to ask about it. One is about infrastructure: is it up, how fast is it, how many requests fail? The other is about content: what did users actually send, and what did the model say back? Model Serving answers the first with endpoint health metrics, which appear by default in the Serving UI. The second needs the payloads themselves, and that is the job of AI Gateway-enabled inference tables.
| Tool | What it captures | Typical use | Access |
|---|---|---|---|
| Endpoint health metrics | Latency, request rate, error rate, CPU usage, memory usage | Understanding the performance and health of the serving infrastructure | Shown by default in the Serving UI for the last 14 days |
| AI Gateway-enabled inference tables | Online prediction requests and responses, written to Delta tables managed by Unity Catalog | Monitoring and debugging model quality or responses, building training datasets, compliance audits | Enabled on new or existing endpoints through the Serving UI or REST API |
Since the log is a Delta table, you can use it for more than reading. Databricks lists several common uses. You can join it with ground truth labels to build a training corpus, and use Lakeflow Jobs to automate retraining from it. You can point data profiling at it to get dashboards and alerts on data drift and model quality. You can debug production issues, because the table records HTTP status codes, the request and response JSON, model run times and traces. And you can use it to monitor deployed agents. Keep in mind that inference tables are a billed feature: Databricks charges for the requests and responses they log.
Checkpoint 1 of 7· Check yourself
A team wants a dataset for fine-tuning their served model. Which inference-table workflow does Databricks describe for this?
The request/response payloads in the inference table, joined with ground truth labels, become a training corpus. Health metrics hold only infrastructure numbers, and you cannot point an inference table at an existing table.
“By joining inference tables with ground truth labels, you can create a training corpus”Source: docs.databricks.com
2.Enabling inference tables and the permissions it needs
AI Gateway-enabled inference tables work on endpoints that serve provisioned throughput workloads, pay-per-token models, external models, deployed agents and custom models. In the Serving UI, you turn them on in the endpoint's AI Gateway section, either when you create the endpoint or later through Edit AI Gateway. To turn them off, you go back to Edit AI Gateway, clear the checkbox and click Update. The newer Unity Gateway experience has its own flow, which starts from AI Gateway in the sidebar. There you can only set up the inference table after the model service already exists.
Checkpoint 2 of 7· Put it in order
Put the steps for enabling an inference table on an existing Unity Gateway model service in order.
- 1.Click Save
- 2.Specify the catalog and schema where the inference table will be stored
- 3.In the sidebar, click AI Gateway
- 4.Click the model service name to open the model service page
- 5.Click Set Up next to Inference tables
You go to the model service first, then click Set Up next to Inference tables, pick the catalog and schema, and save. You never name a table, because Databricks creates it for you.
“Click Set Up next to Inference tables.”Source: docs.databricks.com
Notice that you choose a catalog and schema but never a table name. You cannot point logging at a table that already exists. Databricks creates the table itself. The permissions are checked against both whoever created the endpoint and whoever later changes it. Each needs Can Manage on the endpoint, plus USE CATALOG on the catalog, USE SCHEMA on the schema and CREATE TABLE in the schema. The workspace must have Unity Catalog enabled, and the target catalog cannot be an OpenSharing catalog. The user who enabled the table owns it, and its ACLs follow standard Unity Catalog permissions.
Both of yours. The creator and the modifier each need Can Manage on the endpoint, and each needs USE CATALOG, USE SCHEMA and CREATE TABLE on the target location.
Checkpoint 3 of 7· Exam question
A team deploys an agent to a Model Serving endpoint using the `mlflow.deploy()` API and wants to inspect the exact raw JSON requests and responses that production users have sent to the endpoint over the past week. Where should they look, and how should they query it?
Correct answer: A — Query the Delta table in Unity Catalog that the endpoint's inference table writes to, since `mlflow.deploy()` automatically enables payload logging for the agent
- A. Correct. Deploying an agent with `mlflow.deploy()` automatically enables inference tables, which log raw request and response JSON payloads as rows in a Delta table stored in Unity Catalog, queryable with standard SQL.
- B. Incorrect. MLflow experiment runs and tags are used for tracking evaluation runs and traces, not for storing the raw serving payloads captured for a live endpoint, which are written to a Delta table instead.
- C. Incorrect. Model Serving does not rely on a rotating log file for payload capture; captured requests and responses are persisted durably as Delta table rows in Unity Catalog, not ephemeral logs.
- D. Incorrect. AI Gateway Usage Tables track consumption metrics such as token counts and request counts for governance and cost purposes; they do not store the raw request and response payload content.
3.Reading the schema: latency, status and who called
Every column you need to monitor an endpoint is a first-class field, so you rarely have to parse the JSON payloads for basic health questions. The table below shows the schema for inference tables on serving endpoints. The column to read most carefully is execution_duration_ms. It measures only the time the model spent generating predictions, and it excludes network overhead.
| Column | Type | What it tells you |
|---|---|---|
| databricks_request_id | STRING | Databricks-generated identifier attached to every model serving request |
| client_request_id | STRING | User-provided identifier that can be set in the request body |
| request_time | TIMESTAMP | When the request was received |
| status_code | INT | HTTP status code returned from the model |
| sampling_fraction | DOUBLE | Between 0 and 1; 1 means 100% of requests were included |
| execution_duration_ms | BIGINT | Model inference time only, excluding network latency overhead |
| request / response | STRING | Raw JSON bodies sent to and returned by the endpoint |
| served_entity_id | STRING | Unique ID of the served entity |
| logging_error_codes | ARRAY | Why a row could not be logged, such as MAX_REQUEST_SIZE_EXCEEDED |
| requester | STRING | User or service principal whose permissions were used; NULL for route-optimized custom model endpoints |
If you need details of the foundation model behind an endpoint, join on served_entity_id to the system.serving.served_entities system table.
SELECT * FROM <catalog>.<schema>.<payload_table> payload
JOIN system.serving.served_entities se on payload.served_entity_id = se.served_entity_idInference tables for Unity Gateway model services use different column names. They have request_id, event_time, latency_ms (total latency) and time_to_first_byte_ms, along with destination_type, destination_name and destination_model. They also add invocation_id, because one request can trigger several inference calls, for example a guardrail check followed by the model call.
Checkpoint 4 of 7· Check yourself
Users say a serving endpoint feels slow, but the median execution_duration_ms in its inference table looks fine. What is the most likely explanation?
The column measures only how long the model took to generate predictions. Network latency that users feel is not included.
“This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”Source: docs.databricks.com
4.Where the log has gaps
An inference table is not a complete, real-time record, and it helps to know where the gaps are before you build dashboards on it. Logs are delivered through Zerobus, so they share its latency, its at-least-once delivery (duplicate rows are possible) and its retention behaviour. Requests or responses larger than 10 MiB are not logged; instead, logging_error_codes records MAX_REQUEST_SIZE_EXCEEDED or MAX_RESPONSE_SIZE_EXCEEDED. Requests that return 401, 403, 429 or 500 errors may not be logged at all. That matters when you investigate rate limiting or outages, because the table may under-report exactly those failures. Inference tables also cannot be created in storage secured through a private endpoint.
The table is also easy to break. Changing its schema, renaming it or deleting it can stop logging or corrupt the table. If you want a reshaped or cleaned-up version, build a view or a downstream table on top of it rather than altering the original.
Checkpoint 5 of 7· Check yourself
An analyst counts rows with status_code = 429 in an inference table to measure how often clients are rate-limited. Why is this count unreliable?
The documentation warns that logs may not be populated for 401, 403, 429 and 500 responses, so counting them in the table undercounts.
“Logs may not be populated for requests that return 401, 403, 429, or 500 errors.”Source: docs.databricks.com
Sources3
5.Inference tables for deployed agents
Agents served on Model Serving get a richer inference table. As well as the payload and request details, it stores MLflow Trace logs. These capture the steps inside each agent run, which makes it much easier to debug a bad answer than the final response alone. How you turn this on depends on how you deploy. Agents deployed with the mlflow.deploy() API have inference tables enabled automatically. For programmatic deployments, you set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.
Checkpoint 6 of 7· Check yourself
You deploy an agent programmatically rather than through the mlflow.deploy() API, and you want its inference table to include MLflow traces. What should you do?
Only mlflow.deploy() turns this on automatically. Programmatic deployments need the ENABLE_MLFLOW_TRACING environment variable.
“For programmatic deployments, set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.”Source: docs.databricks.com
Inference tables are scoped to one endpoint, and each endpoint has its own table. For new deployments, Databricks now recommends the unified trace table, which logs all Unity Gateway traffic into a single OpenTelemetry-format table. It describes inference tables as designed for request/response monitoring of a single endpoint. Exam questions about tracking one live endpoint's payloads still point to inference tables.
Checkpoint 7 of 7· Exam question
An ML engineer wants continuous quality monitoring for a live agent endpoint without paying the compute cost of scoring every single production interaction. They plan to use MLflow scorers on a recurring schedule against a subset of traffic. Which configuration approach matches this goal?
Correct answer: A — Register the scorers and start them with a `ScorerSamplingConfig` that sets `sample_rate` below 1.0, so only a configurable percentage of new traces are evaluated on each scheduled run
- A. Correct. Registering scorers and starting them with a `ScorerSamplingConfig` whose `sample_rate` is set below 1.0 lets the monitoring job evaluate only a configurable fraction of production traces on each scheduled pass, controlling cost.
- B. Incorrect. Scorer scheduling is independent of endpoint autoscaling; there is no mechanism that ties trace evaluation to scale-up events, and doing so would leave low-traffic periods completely unmonitored.
- C. Incorrect. Rescoring the entire inference table every night evaluates 100% of traffic rather than a sampled subset, which defeats the stated goal of reducing the compute cost of exhaustive scoring.
- D. Incorrect. There is no `batch_mode` flag that runs scorers inline inside the serving container; production monitoring scorers run as a separate scheduled job against logged traces, not synchronously during inference.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.You can point an inference table at an existing Delta table you created in advance.Why is that wrong?
Databricks always creates a new inference table itself. You only choose the catalog and schema.
Covered in Enabling inference tables and the permissions it needs
2.An inference table is a complete log of every request, so you can safely rename or reshape it for reporting.Why is that wrong?
Large payloads and some error responses are not logged. Changing the table's schema or name, or deleting it, can stop logging or corrupt the table.
Covered in Where the log has gaps
3.execution_duration_ms is the end-to-end latency the user experiences.Why is that wrong?
It measures only model inference time and leaves out network overhead.
Covered in Reading the schema: latency, status and who called
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Inference tables log data like HTTP status codes, request and response JSON code, model run times, and traces output during model run times.”
↩︎ What an inference table gives you that health metrics do not“Provisioned throughput workloads Pay-per-token models External models Deployed agent Custom models”
↩︎ Enabling inference tables and the permissions it needs“In the AI Gateway section, select Enable inference tables.”
↩︎ Enabling inference tables and the permissions it needs“Both the creator of the endpoint and the modifier must have Can Manage permission on the endpoint.”
↩︎ Enabling inference tables and the permissions it needs“Foundation model details are captured in the system.serving.served_entities system table.”
↩︎ Reading the schema: latency, status and who called“You can also enable inference tables for deployed agents, these inference tables store payload and request details as well as MLflow Trace logs.”
↩︎ Inference tables for deployed agents“Agents deployed using the mlflow.deploy() API have inference tables automatically enabled.”
↩︎ Inference tables for deployed agents“The inference table automatically captures incoming requests and outgoing responses for an endpoint and logs them as a Unity Catalog Delta table.”
↩︎ Key concept“Specifying an existing table is not supported.”
↩︎ Exam trap 1“This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”
↩︎ Exam trap 3“By joining inference tables with ground truth labels, you can create a training corpus”
↩︎ Checkpoint“This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”
↩︎ Checkpoint“For programmatic deployments, set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/machine-learning/model-serving/monitor-diagnose-endpointsOfficial docs
“Provides insights into infrastructure metrics like latency, request rate, error rate, CPU usage, and memory usage.”
↩︎ What an inference table gives you that health metrics do not - 3.
“Inference tables are a billed Unity Gateway feature.”
↩︎ What an inference table gives you that health metrics do not“Inference tables can only be configured after you create a model service.”
↩︎ Enabling inference tables and the permissions it needs“Multiple invocations can share the same request_id, such as guardrail checks or multi-turn agent calls.”
↩︎ Reading the schema: latency, status and who called“Inference tables deliver logs through Zerobus and share its latency, at-least-once delivery, and retention behavior.”
↩︎ Where the log has gaps“Requests and responses larger than 10 MiB aren't logged.”
↩︎ Where the log has gaps“The inference table could stop logging data or become corrupted if you do any of the following:”
↩︎ Exam trap 2“Click Set Up next to Inference tables.”
↩︎ Checkpoint“Rows can take up to an hour to appear in the table after the endpoint receives its first request.”
↩︎ Prediction“Logs may not be populated for requests that return 401, 403, 429, or 500 errors.”
↩︎ Checkpoint - 4.
“Inference tables remain available but are designed for single model-service endpoint, request/response monitoring only.”
↩︎ Inference tables for deployed agents