CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 50/56

    Inference Tables for Deployed RAG Apps: What Gets Logged

    Use inference logging to assess deployed RAG application performance

    7 min read
    1.79% of exam
    2 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why a deployed RAG application needs production logging that captures intermediate steps, not just inputs and outputs
    • Enable inference tables on a Model Serving endpoint or a deployed agent
    • Read the inference table schema and say which columns answer which performance question

    Key concept

    Inference table — A Unity Catalog Delta table that Databricks creates and fills automatically with every request and response sent to a serving endpoint. For deployed agents it also stores the MLflow trace. It turns live RAG traffic into data you can query.

    1.Why a deployed RAG app needs its own logs

    Evaluation and monitoring answer the same question, whether the application is good enough, at different times. Databricks puts it simply: evaluation happens during development, and monitoring happens once the app is deployed to production. During evaluation you choose the inputs and run them through a curated evaluation set. In production your users choose the inputs, there are far more of them, and nobody has written the expected answers ahead of time.

    A RAG application is a compound system, so a low-quality answer can come from more than one place. The Databricks RAG cookbook names the diagnostic question production logging has to answer: you need to know if the root cause of a low-quality answer is due to the retrieval step or a hallucination. That is why logging only the final response isn't enough. The log has to keep enough of each request's path to rebuild how the answer was produced.

    On Databricks, the mechanism that captures live traffic is the inference table. For agents it captures the MLflow trace as well as the payload. The rest of this page covers how to turn it on and what it records.

    Checkpoint 1 of 5· Check yourself

    According to the Databricks RAG cookbook, what separates production trace logging from the evaluation harness used in development?

    Sources1

    2.Turning on inference tables for an endpoint or agent

    AI Gateway-enabled inference tables work on endpoints that serve provisioned throughput workloads, pay-per-token models, external models, deployed agents, and custom models. You can enable them on a new endpoint or an existing one. Either way, every request to the endpoint from then on is logged automatically to a table in Unity Catalog.

    From the Serving UI, you enable them while creating the endpoint:

    Checkpoint 2 of 5· Put it in order

    Put these steps for enabling inference tables during endpoint creation in order.

    1. 1.Click Serving in the Databricks UI
    2. 2.In the AI Gateway section, select Enable inference tables
    3. 3.Click Create serving endpoint

    Most RAG applications on Databricks are deployed as agents, and that is the case that matters for this objective. For agents, the inference table stores payload and request details as well as MLflow Trace logs. That is how the retrieval step from the previous section ends up in the log. There are two ways to get it:

    - Agents deployed with the mlflow.deploy() API have inference tables turned on automatically. - For programmatic deployments, set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.

    You also need the right permissions. Both the creator and the modifier of the endpoint need Can Manage on the endpoint, plus USE CATALOG, USE SCHEMA, and CREATE TABLE in the target catalog and schema. You can't point the feature at an existing table. Databricks creates a new one for you.

    Checkpoint 3 of 5· Exam question

    An ML engineer wants to build a new fine-tuning and evaluation dataset by combining real production traffic with SME-provided correct answers. They plan to join AI Gateway inference table records against a table of human-labeled ground truth. Which pair of inference table columns provides the raw content needed for this join?

    Checkpoint 4 of 5· Check yourself

    You deploy a RAG agent programmatically, not with the agent deployment API, and the inference table has no traces. What is the documented fix?

    Sources2

    3.Reading the schema: which column answers which question

    To assess performance from the log, you need to know what each column actually measures. The AI Gateway-enabled inference table for a serving endpoint has one row per request. Here are the columns you will use most:

    AI Gateway-enabled inference table for serving endpoints: key columns and the performance question each one answers
    ColumnWhat it holdsUse it to assess
    databricks_request_idDatabricks-generated identifier attached to every requestTracing one specific request end to end
    request_date / request_timeUTC date and timestamp when the request was receivedTrends over time, partitioning queries
    status_codeHTTP status code returned from the modelError rates
    execution_duration_msTime the model spent on inference, excluding network overheadModel compute time (not end-to-end latency)
    request / responseRaw JSON bodies sent and returnedWhat users asked and what the app answered
    sampling_fractionBetween 0 and 1; 1 means 100% of requests were includedWhether counts need scaling up
    served_entity_idUnique ID of the served entityJoining to model details
    requesterUser or service principal whose permissions were usedWho is calling the app
    logging_error_codesErrors when data could not be logged, e.g. MAX_REQUEST_SIZE_EXCEEDEDGaps in the log itself

    Two columns are easy to misread. execution_duration_ms covers only the model's own inference time. Network overhead isn't included, so it understates what the user actually waits. requester returns NULL for route-optimized custom model endpoints, so per-user analysis won't work on those.

    Deployed agents get a payload request logs table built for conversations. As well as raw request and response strings, it pulls out fields that make RAG analysis much easier:

    Agent payload request logs table: conversation-aware columns
    ColumnDescription
    conversation_idThe conversation id extracted from request logs
    requestThe last user query from the user's conversation
    responseThe last response to the user
    request_raw / response_rawString representation of the request / response
    execution_time_msTime the model performed inference, excluding network overhead
    traceString representation of the trace

    Checkpoint 5 of 5· Match them up

    Match each RAG performance question to the inference table column that answers it.

    Tap a term, then the definition that fits it.

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.execution_duration_ms in a serving-endpoint inference table is the end-to-end latency the user experiences.Why is that wrong?

      It measures only the model's inference time. Network overhead is excluded.

      Covered in Reading the schema: which column answers which question

    2. 2.You can send inference logs to a Delta table you have already created, for example one that holds your evaluation set.Why is that wrong?

      Databricks always creates a new inference table itself. Pointing at an existing table is not supported.

      Covered in Turning on inference tables for an endpoint or agent

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “evaluation happens during development and monitoring happens once the application is deployed to production”
      ↩︎ Why a deployed RAG app needs its own logs
      “you need to know if the root cause of a low-quality answer is due to the retrieval step or a hallucination”
      ↩︎ Why a deployed RAG app needs its own logs
      “Your production logging must track the inputs, outputs, and intermediate steps such as document retrieval”
      ↩︎ Prediction
      “Once in production, you need to evaluate a significantly higher quantity of requests/responses and how each response was generated.”
      ↩︎ Checkpoint
    2. 2.
      “Provisioned throughput workloads Pay-per-token models External models Deployed agent Custom models”
      ↩︎ Turning on inference tables for an endpoint or agent
      “these inference tables store payload and request details as well as MLflow Trace logs.”
      ↩︎ Turning on inference tables for an endpoint or agent
      “Agents deployed using the mlflow.deploy() API have inference tables automatically enabled.”
      ↩︎ Turning on inference tables for an endpoint or agent
      “This field returns NULL for route-optimized custom model endpoints.”
      ↩︎ Reading the schema: which column answers which question
      “The last user query from the user's conversation.”
      ↩︎ Reading the schema: which column answers which question
      “The inference table automatically captures incoming requests and outgoing responses for an endpoint and logs them as a Unity Catalog Delta table.”
      ↩︎ Key concept
      “This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”
      ↩︎ Exam trap 1
      “Specifying an existing table is not supported.”
      ↩︎ Exam trap 2
      “In the AI Gateway section, select Enable inference tables.”
      ↩︎ Checkpoint
      “For programmatic deployments, set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.”
      ↩︎ Checkpoint
      “This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Querying Inference Logs to Diagnose Deployed RAG Performance

    Spotted a mistake, or was something unclear? Tell us.