What you will be able to do
- Query an inference table with SQL and join it to served-model details
- Pick the right use of logged data: debugging, quality and drift monitoring, building a training corpus, or comparing performance
- Account for the gaps in the logs before drawing conclusions
- Explain how logged traces feed scheduled MLflow scorers and how inference tables differ from the unified trace table
1.Querying the log
An inference table is an ordinary Unity Catalog Delta table. Once the served models are ready, every request and its response are logged automatically. You can open the table in Catalog Explorer from the endpoint page, query it from Databricks SQL or a notebook, or query it through the REST API. The simplest place to start is a full scan:
SELECT * FROM <catalog>.<schema>.<payload_table>The payload itself doesn't say which foundation model produced each answer. That detail lives in the system.serving.served_entities system table, and you join to it on served_entity_id. When you are comparing generator models behind the same RAG endpoint, this join is what lets you group quality or latency by model:
SELECT * FROM <catalog>.<schema>.<payload_table> payload
JOIN system.serving.served_entities se on payload.served_entity_id = se.served_entity_idCheckpoint 1 of 5· Check yourself
You want to break down your RAG endpoint's logged requests by the underlying foundation model. Which table do you join the inference table to?
Foundation model details sit in the served_entities system table, and the documented join key is served_entity_id.
“Foundation model details are captured in the system.serving.served_entities system table.”Source: docs.databricks.com
Sources1
2.Four ways to use the logged data
Databricks names four common uses. Each one answers a different performance question:
| Application | How it works | What it tells you about a deployed RAG app |
|---|---|---|
| Debug production issues | Logs HTTP status codes, request/response JSON, model run times, and traces | Why a specific request failed or ran slowly |
| Monitor data and model quality | Data profiling generates data and model quality dashboards; alerts fire on shifts | Whether incoming questions are drifting or quality is dropping |
| Create a training corpus | Join inference tables with ground truth labels; automate retraining with Lakeflow Jobs | Real user traffic to fine-tune or improve the model |
| Monitor deployed agents | Inference tables store MLflow traces for agents | Step-level view of retrieval and generation |
History matters too. Because the log keeps past traffic, you can use it to compare model performance on historical requests, for example by replaying last month's real questions against a candidate configuration. For a RAG app, the agent-trace use is the one that connects back to root cause. The trace records the retrieval step, so a bad answer can be traced either to the documents fetched or to what the generator did with them.
Checkpoint 2 of 5· Exam question
A production monitoring job for a deployed RAG agent runs an expensive LLM-as-judge scorer that checks groundedness on every incoming trace, and the team notices monitoring costs rising sharply as traffic grows. Which change addresses the cost concern while still giving ongoing visibility into groundedness quality?
Correct answer: A — Lower the scorer's sampling rate so groundedness is evaluated on a representative subset of traces instead of all traffic
- A. Reducing the sampling rate lets a scheduled scorer evaluate a representative fraction of live traces rather than every single one, which is the documented way to balance evaluation coverage against the cost of running an expensive LLM judge continuously.
- B. HTTP status codes only indicate whether a call completed successfully, not whether the response was grounded in retrieved context, so dropping the scorer entirely would eliminate the quality signal the team needs.
- C. Restricting groundedness checks to offline evaluation before deployment removes the ability to detect quality drift on real production traffic, which defeats the purpose of production monitoring.
- D. The payload size limit controls whether an individual request or response body is captured, not how many traces a scheduled scorer processes, so raising it would not reduce scoring cost or address the sampling tradeoff.
Checkpoint 3 of 5· Check yourself
A team wants to retrain a model on real production questions paired with correct answers that reviewers wrote later. What does the documentation recommend?
Inference tables hold requests and responses but no labels. You add the labels with a join, and Lakeflow Jobs can automate the retraining loop.
“By joining inference tables with ground truth labels, you can create a training corpus”Source: docs.databricks.com
Sources1
3.What the log can miss
A performance number pulled from logs is only as good as the logs' coverage. The Unity Gateway inference-table documentation lists several limits:
- Error responses. Requests that return 401, 403, 429, or 500 may not be logged, so an error rate built from status_code can be too low.
- Payload size. Requests and responses larger than 10 MiB aren't logged. When that happens, logging_error_codes shows MAX_REQUEST_SIZE_EXCEEDED or MAX_RESPONSE_SIZE_EXCEEDED. Long RAG contexts make this a real risk.
- Delay. Rows can take up to an hour to appear after the endpoint receives its first request.
- Sampling. sampling_fraction records down-sampling, and a value of 1 means none. Below 1, raw row counts understate real volume.
You can also break logging yourself. The inference table could stop logging data or become corrupted if you change its schema, rename it, or delete it. Build your analysis as views or downstream tables, and leave the inference table itself alone.
Checkpoint 4 of 5· Check yourself
An analyst adds a derived column to the inference table so that a dashboard can read it directly. What is the risk?
Changing the schema, changing the name, or deleting the table can stop logging or corrupt the table.
“The inference table could stop logging data or become corrupted if you do any of the following:”Source: docs.databricks.com
Sources2
4.From raw logs to scored quality, and the unified trace table
Inference tables tell you what happened. They don't tell you whether an answer was any good. For quality, Databricks offers production monitoring, which lets you automatically run MLflow 3 scorers on traces from your agents to continuously assess quality. One prerequisite links it to everything above: the agent must log traces using MLflow Tracing. A scorer is registered against the experiment and then started with a sampling rate. The results are attached as feedback on each trace it evaluates. The docs suggest sample_rate=1.0 for critical checks such as safety, and 0.05–0.2 for expensive LLM judges.
Checkpoint 5 of 5· Fill the gap
Which method begins monitoring live traces once the scorer is registered?
from mlflow.genai.scorers import Safety, ScorerSamplingConfig
# Register and start a built-in judge
safety_judge = Safety().register(name="safety")
safety_judge = safety_judge. ? (sampling_config=ScorerSamplingConfig(sample_rate=0.7))Production monitoring follows a two-step .register() then .start() pattern. Monitoring begins the moment the scorer is started.
Source: docs.databricks.comThere is also a newer logging option. The unified trace table, currently in Beta, captures every request and response across all Unity Gateway services in one table in OpenTelemetry format. Databricks calls it the recommended approach for new deployments, and says inference tables are designed for single-endpoint request/response monitoring only. The difference shows most clearly with multi-hop agents:
| Dimension | Unified trace table | Inference tables |
|---|---|---|
| Scope | All Unity Gateway model services and MCP services, in one table | Per model serving endpoint, one table each |
| Setup | One-time setup by metastore admin | Must be enabled per endpoint |
| Schema | OpenTelemetry spans | Databricks-specific (requires post-processing) |
| Agentic workflows | All hops land in one table | Fragmented across per-endpoint tables |
| MLflow compatibility | MLflow-compatible OTel schema that MLflow tools can read directly | Requires manual trace extraction |
| Owner | Metastore admin who creates the table | Endpoint owner |
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Counting non-200 rows in the inference table gives the complete error rate for a deployed RAG endpoint.Why is that wrong?
Some error responses (401, 403, 429, 500) may not be logged at all, so a rate built from logged rows can understate failures.
Covered in What the log can miss
2.Inference tables are the recommended way to log multi-hop agent traffic across many Unity Gateway services.Why is that wrong?
Inference tables are per endpoint and fragment agentic workflows. Databricks recommends the unified trace table for new deployments.
Covered in From raw logs to scored quality, and the unified trace table
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“You can view the table in the UI, query the table from Databricks SQL or a notebook, or query the table using the REST API.”
↩︎ Querying the log“Inference tables log data like HTTP status codes, request and response JSON code, model run times, and traces output during model run times.”
↩︎ Four ways to use the logged data“You can continuously monitor your model performance and data drift using data profiling”
↩︎ Four ways to use the logged data“You can also use the historical data in inference tables to compare model performance on historical requests.”
↩︎ Four ways to use the logged data“Foundation model details are captured in the system.serving.served_entities system table.”
↩︎ Checkpoint“By joining inference tables with ground truth labels, you can create a training corpus”
↩︎ Checkpoint“The inference table could stop logging data or become corrupted if you do any of the following:”
↩︎ Checkpoint - 2.
“Requests and responses larger than 10 MiB aren't logged.”
↩︎ What the log can miss“Rows can take up to an hour to appear in the table after the endpoint receives its first request.”
↩︎ What the log can miss“Logs may not be populated for requests that return 401, 403, 429, or 500 errors.”
↩︎ Exam trap 1“Logs may not be populated for requests that return 401, 403, 429, or 500 errors.”
↩︎ Prediction - 3.
“Production monitoring lets you automatically run MLflow 3 scorers on traces from your agents to continuously assess quality.”
↩︎ From raw logs to scored quality, and the unified trace table“For expensive scorers, such as complex LLM judges, use lower sample rates (0.05-0.2).”
↩︎ From raw logs to scored quality, and the unified trace table - 4.
“The unified trace table is the recommended approach for new deployments.”
↩︎ From raw logs to scored quality, and the unified trace table“Inference tables remain available but are designed for single model-service endpoint, request/response monitoring only.”
↩︎ Exam trap 2