CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 52/56

    Inference Tables: Logging a Live LLM or Agent Endpoint

    Use inference tables and Agent Monitoring to track a live LLM endpoint

    11 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what an AI Gateway-enabled inference table captures and what it is used for
    • Enable inference tables on a serving endpoint and name the permissions this requires
    • Read the inference table schema to work out latency, status and requester for each request
    • Spot the limits that leave gaps in an inference table: payload size, error codes, delivery delay and table changes
    • Turn on inference tables with MLflow traces for a deployed agent

    Key concept

    Inference table — An inference table is a Unity Catalog Delta table. Databricks fills it automatically with every request a serving endpoint receives and every response it returns. Because it is ordinary governed data, you can query it with SQL to monitor, debug, audit and improve a live model or agent.

    1.What an inference table gives you that health metrics do not

    Once an LLM endpoint is live, there are two kinds of question to ask about it. One is about infrastructure: is it up, how fast is it, how many requests fail? The other is about content: what did users actually send, and what did the model say back? Model Serving answers the first with endpoint health metrics, which appear by default in the Serving UI. The second needs the payloads themselves, and that is the job of AI Gateway-enabled inference tables.

    Two Model Serving monitoring tools and the questions each one answers
    ToolWhat it capturesTypical useAccess
    Endpoint health metricsLatency, request rate, error rate, CPU usage, memory usageUnderstanding the performance and health of the serving infrastructureShown by default in the Serving UI for the last 14 days
    AI Gateway-enabled inference tablesOnline prediction requests and responses, written to Delta tables managed by Unity CatalogMonitoring and debugging model quality or responses, building training datasets, compliance auditsEnabled on new or existing endpoints through the Serving UI or REST API

    Since the log is a Delta table, you can use it for more than reading. Databricks lists several common uses. You can join it with ground truth labels to build a training corpus, and use Lakeflow Jobs to automate retraining from it. You can point data profiling at it to get dashboards and alerts on data drift and model quality. You can debug production issues, because the table records HTTP status codes, the request and response JSON, model run times and traces. And you can use it to monitor deployed agents. Keep in mind that inference tables are a billed feature: Databricks charges for the requests and responses they log.

    Checkpoint 1 of 7· Check yourself

    A team wants a dataset for fine-tuning their served model. Which inference-table workflow does Databricks describe for this?

    Sources123

    2.Enabling inference tables and the permissions it needs

    AI Gateway-enabled inference tables work on endpoints that serve provisioned throughput workloads, pay-per-token models, external models, deployed agents and custom models. In the Serving UI, you turn them on in the endpoint's AI Gateway section, either when you create the endpoint or later through Edit AI Gateway. To turn them off, you go back to Edit AI Gateway, clear the checkbox and click Update. The newer Unity Gateway experience has its own flow, which starts from AI Gateway in the sidebar. There you can only set up the inference table after the model service already exists.

    Checkpoint 2 of 7· Put it in order

    Put the steps for enabling an inference table on an existing Unity Gateway model service in order.

    1. 1.Click Save
    2. 2.Specify the catalog and schema where the inference table will be stored
    3. 3.In the sidebar, click AI Gateway
    4. 4.Click the model service name to open the model service page
    5. 5.Click Set Up next to Inference tables

    Notice that you choose a catalog and schema but never a table name. You cannot point logging at a table that already exists. Databricks creates the table itself. The permissions are checked against both whoever created the endpoint and whoever later changes it. Each needs Can Manage on the endpoint, plus USE CATALOG on the catalog, USE SCHEMA on the schema and CREATE TABLE in the schema. The workspace must have Unity Catalog enabled, and the target catalog cannot be an OpenSharing catalog. The user who enabled the table owns it, and its ACLs follow standard Unity Catalog permissions.

    Checkpoint 3 of 7· Exam question

    A team deploys an agent to a Model Serving endpoint using the `mlflow.deploy()` API and wants to inspect the exact raw JSON requests and responses that production users have sent to the endpoint over the past week. Where should they look, and how should they query it?

    Sources13

    3.Reading the schema: latency, status and who called

    Every column you need to monitor an endpoint is a first-class field, so you rarely have to parse the JSON payloads for basic health questions. The table below shows the schema for inference tables on serving endpoints. The column to read most carefully is execution_duration_ms. It measures only the time the model spent generating predictions, and it excludes network overhead.

    Selected columns of an AI Gateway-enabled inference table on a serving endpoint
    ColumnTypeWhat it tells you
    databricks_request_idSTRINGDatabricks-generated identifier attached to every model serving request
    client_request_idSTRINGUser-provided identifier that can be set in the request body
    request_timeTIMESTAMPWhen the request was received
    status_codeINTHTTP status code returned from the model
    sampling_fractionDOUBLEBetween 0 and 1; 1 means 100% of requests were included
    execution_duration_msBIGINTModel inference time only, excluding network latency overhead
    request / responseSTRINGRaw JSON bodies sent to and returned by the endpoint
    served_entity_idSTRINGUnique ID of the served entity
    logging_error_codesARRAYWhy a row could not be logged, such as MAX_REQUEST_SIZE_EXCEEDED
    requesterSTRINGUser or service principal whose permissions were used; NULL for route-optimized custom model endpoints

    If you need details of the foundation model behind an endpoint, join on served_entity_id to the system.serving.served_entities system table.

    Joining inference-table rows to details of the served foundation modelsql
    SELECT * FROM <catalog>.<schema>.<payload_table> payload
    JOIN system.serving.served_entities se on payload.served_entity_id = se.served_entity_id

    Inference tables for Unity Gateway model services use different column names. They have request_id, event_time, latency_ms (total latency) and time_to_first_byte_ms, along with destination_type, destination_name and destination_model. They also add invocation_id, because one request can trigger several inference calls, for example a guardrail check followed by the model call.

    Checkpoint 4 of 7· Check yourself

    Users say a serving endpoint feels slow, but the median execution_duration_ms in its inference table looks fine. What is the most likely explanation?

    Sources13

    4.Where the log has gaps

    An inference table is not a complete, real-time record, and it helps to know where the gaps are before you build dashboards on it. Logs are delivered through Zerobus, so they share its latency, its at-least-once delivery (duplicate rows are possible) and its retention behaviour. Requests or responses larger than 10 MiB are not logged; instead, logging_error_codes records MAX_REQUEST_SIZE_EXCEEDED or MAX_RESPONSE_SIZE_EXCEEDED. Requests that return 401, 403, 429 or 500 errors may not be logged at all. That matters when you investigate rate limiting or outages, because the table may under-report exactly those failures. Inference tables also cannot be created in storage secured through a private endpoint.

    The table is also easy to break. Changing its schema, renaming it or deleting it can stop logging or corrupt the table. If you want a reshaped or cleaned-up version, build a view or a downstream table on top of it rather than altering the original.

    Checkpoint 5 of 7· Check yourself

    An analyst counts rows with status_code = 429 in an inference table to measure how often clients are rate-limited. Why is this count unreliable?

    Sources3

    5.Inference tables for deployed agents

    Agents served on Model Serving get a richer inference table. As well as the payload and request details, it stores MLflow Trace logs. These capture the steps inside each agent run, which makes it much easier to debug a bad answer than the final response alone. How you turn this on depends on how you deploy. Agents deployed with the mlflow.deploy() API have inference tables enabled automatically. For programmatic deployments, you set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.

    Checkpoint 6 of 7· Check yourself

    You deploy an agent programmatically rather than through the mlflow.deploy() API, and you want its inference table to include MLflow traces. What should you do?

    Inference tables are scoped to one endpoint, and each endpoint has its own table. For new deployments, Databricks now recommends the unified trace table, which logs all Unity Gateway traffic into a single OpenTelemetry-format table. It describes inference tables as designed for request/response monitoring of a single endpoint. Exam questions about tracking one live endpoint's payloads still point to inference tables.

    Checkpoint 7 of 7· Exam question

    An ML engineer wants continuous quality monitoring for a live agent endpoint without paying the compute cost of scoring every single production interaction. They plan to use MLflow scorers on a recurring schedule against a subset of traffic. Which configuration approach matches this goal?

    Sources14

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.You can point an inference table at an existing Delta table you created in advance.Why is that wrong?

      Databricks always creates a new inference table itself. You only choose the catalog and schema.

      Covered in Enabling inference tables and the permissions it needs

    2. 2.An inference table is a complete log of every request, so you can safely rename or reshape it for reporting.Why is that wrong?

      Large payloads and some error responses are not logged. Changing the table's schema or name, or deleting it, can stop logging or corrupt the table.

      Covered in Where the log has gaps

    3. 3.execution_duration_ms is the end-to-end latency the user experiences.Why is that wrong?

      It measures only model inference time and leaves out network overhead.

      Covered in Reading the schema: latency, status and who called

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Inference tables log data like HTTP status codes, request and response JSON code, model run times, and traces output during model run times.”
      ↩︎ What an inference table gives you that health metrics do not
      “Provisioned throughput workloads Pay-per-token models External models Deployed agent Custom models”
      ↩︎ Enabling inference tables and the permissions it needs
      “In the AI Gateway section, select Enable inference tables.”
      ↩︎ Enabling inference tables and the permissions it needs
      “Both the creator of the endpoint and the modifier must have Can Manage permission on the endpoint.”
      ↩︎ Enabling inference tables and the permissions it needs
      “Foundation model details are captured in the system.serving.served_entities system table.”
      ↩︎ Reading the schema: latency, status and who called
      “You can also enable inference tables for deployed agents, these inference tables store payload and request details as well as MLflow Trace logs.”
      ↩︎ Inference tables for deployed agents
      “Agents deployed using the mlflow.deploy() API have inference tables automatically enabled.”
      ↩︎ Inference tables for deployed agents
      “The inference table automatically captures incoming requests and outgoing responses for an endpoint and logs them as a Unity Catalog Delta table.”
      ↩︎ Key concept
      “Specifying an existing table is not supported.”
      ↩︎ Exam trap 1
      “This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”
      ↩︎ Exam trap 3
      “By joining inference tables with ground truth labels, you can create a training corpus”
      ↩︎ Checkpoint
      “This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”
      ↩︎ Checkpoint
      “For programmatic deployments, set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.”
      ↩︎ Checkpoint
    2. 2.
      “Provides insights into infrastructure metrics like latency, request rate, error rate, CPU usage, and memory usage.”
      ↩︎ What an inference table gives you that health metrics do not
    3. 3.
      “Inference tables are a billed Unity Gateway feature.”
      ↩︎ What an inference table gives you that health metrics do not
      “Inference tables can only be configured after you create a model service.”
      ↩︎ Enabling inference tables and the permissions it needs
      “Multiple invocations can share the same request_id, such as guardrail checks or multi-turn agent calls.”
      ↩︎ Reading the schema: latency, status and who called
      “Inference tables deliver logs through Zerobus and share its latency, at-least-once delivery, and retention behavior.”
      ↩︎ Where the log has gaps
      “Requests and responses larger than 10 MiB aren't logged.”
      ↩︎ Where the log has gaps
      “The inference table could stop logging data or become corrupted if you do any of the following:”
      ↩︎ Exam trap 2
      “Click Set Up next to Inference tables.”
      ↩︎ Checkpoint
      “Rows can take up to an hour to appear in the table after the endpoint receives its first request.”
      ↩︎ Prediction
      “Logs may not be populated for requests that return 401, 403, 429, or 500 errors.”
      ↩︎ Checkpoint
    4. 4.
      “Inference tables remain available but are designed for single model-service endpoint, request/response monitoring only.”
      ↩︎ Inference tables for deployed agents

    Continue to page 2 of 2

    Agent Monitoring: Running Scorers on Live Production Traces

    Spotted a mistake, or was something unclear? Tell us.