CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 50/56

    Querying Inference Logs to Diagnose Deployed RAG Performance

    Use inference logging to assess deployed RAG application performance

    8 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Query an inference table with SQL and join it to served-model details
    • Pick the right use of logged data: debugging, quality and drift monitoring, building a training corpus, or comparing performance
    • Account for the gaps in the logs before drawing conclusions
    • Explain how logged traces feed scheduled MLflow scorers and how inference tables differ from the unified trace table

    1.Querying the log

    An inference table is an ordinary Unity Catalog Delta table. Once the served models are ready, every request and its response are logged automatically. You can open the table in Catalog Explorer from the endpoint page, query it from Databricks SQL or a notebook, or query it through the REST API. The simplest place to start is a full scan:

    Reading every logged request for an endpointsql
    SELECT * FROM <catalog>.<schema>.<payload_table>

    The payload itself doesn't say which foundation model produced each answer. That detail lives in the system.serving.served_entities system table, and you join to it on served_entity_id. When you are comparing generator models behind the same RAG endpoint, this join is what lets you group quality or latency by model:

    Joining logged requests to details of the served foundation modelsql
    SELECT * FROM <catalog>.<schema>.<payload_table> payload
    JOIN system.serving.served_entities se on payload.served_entity_id = se.served_entity_id

    Checkpoint 1 of 5· Check yourself

    You want to break down your RAG endpoint's logged requests by the underlying foundation model. Which table do you join the inference table to?

    Sources1

    2.Four ways to use the logged data

    Databricks names four common uses. Each one answers a different performance question:

    Documented applications of inference tables and what each gives a RAG team
    ApplicationHow it worksWhat it tells you about a deployed RAG app
    Debug production issuesLogs HTTP status codes, request/response JSON, model run times, and tracesWhy a specific request failed or ran slowly
    Monitor data and model qualityData profiling generates data and model quality dashboards; alerts fire on shiftsWhether incoming questions are drifting or quality is dropping
    Create a training corpusJoin inference tables with ground truth labels; automate retraining with Lakeflow JobsReal user traffic to fine-tune or improve the model
    Monitor deployed agentsInference tables store MLflow traces for agentsStep-level view of retrieval and generation

    History matters too. Because the log keeps past traffic, you can use it to compare model performance on historical requests, for example by replaying last month's real questions against a candidate configuration. For a RAG app, the agent-trace use is the one that connects back to root cause. The trace records the retrieval step, so a bad answer can be traced either to the documents fetched or to what the generator did with them.

    Checkpoint 2 of 5· Exam question

    A production monitoring job for a deployed RAG agent runs an expensive LLM-as-judge scorer that checks groundedness on every incoming trace, and the team notices monitoring costs rising sharply as traffic grows. Which change addresses the cost concern while still giving ongoing visibility into groundedness quality?

    Checkpoint 3 of 5· Check yourself

    A team wants to retrain a model on real production questions paired with correct answers that reviewers wrote later. What does the documentation recommend?

    Sources1

    3.What the log can miss

    A performance number pulled from logs is only as good as the logs' coverage. The Unity Gateway inference-table documentation lists several limits:

    - Error responses. Requests that return 401, 403, 429, or 500 may not be logged, so an error rate built from status_code can be too low. - Payload size. Requests and responses larger than 10 MiB aren't logged. When that happens, logging_error_codes shows MAX_REQUEST_SIZE_EXCEEDED or MAX_RESPONSE_SIZE_EXCEEDED. Long RAG contexts make this a real risk. - Delay. Rows can take up to an hour to appear after the endpoint receives its first request. - Sampling. sampling_fraction records down-sampling, and a value of 1 means none. Below 1, raw row counts understate real volume.

    You can also break logging yourself. The inference table could stop logging data or become corrupted if you change its schema, rename it, or delete it. Build your analysis as views or downstream tables, and leave the inference table itself alone.

    Checkpoint 4 of 5· Check yourself

    An analyst adds a derived column to the inference table so that a dashboard can read it directly. What is the risk?

    Sources2

    4.From raw logs to scored quality, and the unified trace table

    Inference tables tell you what happened. They don't tell you whether an answer was any good. For quality, Databricks offers production monitoring, which lets you automatically run MLflow 3 scorers on traces from your agents to continuously assess quality. One prerequisite links it to everything above: the agent must log traces using MLflow Tracing. A scorer is registered against the experiment and then started with a sampling rate. The results are attached as feedback on each trace it evaluates. The docs suggest sample_rate=1.0 for critical checks such as safety, and 0.05–0.2 for expensive LLM judges.

    Checkpoint 5 of 5· Fill the gap

    Which method begins monitoring live traces once the scorer is registered?

    from mlflow.genai.scorers import Safety, ScorerSamplingConfig
    
    # Register and start a built-in judge
    safety_judge = Safety().register(name="safety")
    safety_judge = safety_judge. ? (sampling_config=ScorerSamplingConfig(sample_rate=0.7))

    There is also a newer logging option. The unified trace table, currently in Beta, captures every request and response across all Unity Gateway services in one table in OpenTelemetry format. Databricks calls it the recommended approach for new deployments, and says inference tables are designed for single-endpoint request/response monitoring only. The difference shows most clearly with multi-hop agents:

    Unified trace table compared with inference tables
    DimensionUnified trace tableInference tables
    ScopeAll Unity Gateway model services and MCP services, in one tablePer model serving endpoint, one table each
    SetupOne-time setup by metastore adminMust be enabled per endpoint
    SchemaOpenTelemetry spansDatabricks-specific (requires post-processing)
    Agentic workflowsAll hops land in one tableFragmented across per-endpoint tables
    MLflow compatibilityMLflow-compatible OTel schema that MLflow tools can read directlyRequires manual trace extraction
    OwnerMetastore admin who creates the tableEndpoint owner

    Sources34

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Counting non-200 rows in the inference table gives the complete error rate for a deployed RAG endpoint.Why is that wrong?

      Some error responses (401, 403, 429, 500) may not be logged at all, so a rate built from logged rows can understate failures.

      Covered in What the log can miss

    2. 2.Inference tables are the recommended way to log multi-hop agent traffic across many Unity Gateway services.Why is that wrong?

      Inference tables are per endpoint and fragment agentic workflows. Databricks recommends the unified trace table for new deployments.

      Covered in From raw logs to scored quality, and the unified trace table

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “You can view the table in the UI, query the table from Databricks SQL or a notebook, or query the table using the REST API.”
      ↩︎ Querying the log
      “Inference tables log data like HTTP status codes, request and response JSON code, model run times, and traces output during model run times.”
      ↩︎ Four ways to use the logged data
      “You can continuously monitor your model performance and data drift using data profiling”
      ↩︎ Four ways to use the logged data
      “You can also use the historical data in inference tables to compare model performance on historical requests.”
      ↩︎ Four ways to use the logged data
      “Foundation model details are captured in the system.serving.served_entities system table.”
      ↩︎ Checkpoint
      “By joining inference tables with ground truth labels, you can create a training corpus”
      ↩︎ Checkpoint
      “The inference table could stop logging data or become corrupted if you do any of the following:”
      ↩︎ Checkpoint
    2. 2.
      “Requests and responses larger than 10 MiB aren't logged.”
      ↩︎ What the log can miss
      “Rows can take up to an hour to appear in the table after the endpoint receives its first request.”
      ↩︎ What the log can miss
      “Logs may not be populated for requests that return 401, 403, 429, or 500 errors.”
      ↩︎ Exam trap 1
      “Logs may not be populated for requests that return 401, 403, 429, or 500 errors.”
      ↩︎ Prediction
    3. 3.
      “Production monitoring lets you automatically run MLflow 3 scorers on traces from your agents to continuously assess quality.”
      ↩︎ From raw logs to scored quality, and the unified trace table
      “For expensive scorers, such as complex LLM judges, use lower sample rates (0.05-0.2).”
      ↩︎ From raw logs to scored quality, and the unified trace table
    4. 4.
      “The unified trace table is the recommended approach for new deployments.”
      ↩︎ From raw logs to scored quality, and the unified trace table
      “Inference tables remain available but are designed for single model-service endpoint, request/response monitoring only.”
      ↩︎ Exam trap 2

    Spotted a mistake, or was something unclear? Tell us.