What you will be able to do
- Explain why a deployed RAG application needs production logging that captures intermediate steps, not just inputs and outputs
- Enable inference tables on a Model Serving endpoint or a deployed agent
- Read the inference table schema and say which columns answer which performance question
Key concept
Inference table — A Unity Catalog Delta table that Databricks creates and fills automatically with every request and response sent to a serving endpoint. For deployed agents it also stores the MLflow trace. It turns live RAG traffic into data you can query.
1.Why a deployed RAG app needs its own logs
Evaluation and monitoring answer the same question, whether the application is good enough, at different times. Databricks puts it simply: evaluation happens during development, and monitoring happens once the app is deployed to production. During evaluation you choose the inputs and run them through a curated evaluation set. In production your users choose the inputs, there are far more of them, and nobody has written the expected answers ahead of time.
A RAG application is a compound system, so a low-quality answer can come from more than one place. The Databricks RAG cookbook names the diagnostic question production logging has to answer: you need to know if the root cause of a low-quality answer is due to the retrieval step or a hallucination. That is why logging only the final response isn't enough. The log has to keep enough of each request's path to rebuild how the answer was produced.
On Databricks, the mechanism that captures live traffic is the inference table. For agents it captures the MLflow trace as well as the payload. The rest of this page covers how to turn it on and what it records.
Checkpoint 1 of 5· Check yourself
According to the Databricks RAG cookbook, what separates production trace logging from the evaluation harness used in development?
Once deployed, you assess many more requests than any evaluation set, and you need each one's generation path so you can find the root cause.
“Once in production, you need to evaluate a significantly higher quantity of requests/responses and how each response was generated.”Source: docs.databricks.com
Sources1
2.Turning on inference tables for an endpoint or agent
AI Gateway-enabled inference tables work on endpoints that serve provisioned throughput workloads, pay-per-token models, external models, deployed agents, and custom models. You can enable them on a new endpoint or an existing one. Either way, every request to the endpoint from then on is logged automatically to a table in Unity Catalog.
From the Serving UI, you enable them while creating the endpoint:
Checkpoint 2 of 5· Put it in order
Put these steps for enabling inference tables during endpoint creation in order.
- 1.Click Serving in the Databricks UI
- 2.In the AI Gateway section, select Enable inference tables
- 3.Click Create serving endpoint
Inference tables are an AI Gateway setting on the endpoint, so you set them up inside the AI Gateway section of the endpoint you are creating. For an existing endpoint, you click Edit AI Gateway instead.
“In the AI Gateway section, select Enable inference tables.”Source: docs.databricks.com
Most RAG applications on Databricks are deployed as agents, and that is the case that matters for this objective. For agents, the inference table stores payload and request details as well as MLflow Trace logs. That is how the retrieval step from the previous section ends up in the log. There are two ways to get it:
- Agents deployed with the mlflow.deploy() API have inference tables turned on automatically.
- For programmatic deployments, set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.
You also need the right permissions. Both the creator and the modifier of the endpoint need Can Manage on the endpoint, plus USE CATALOG, USE SCHEMA, and CREATE TABLE in the target catalog and schema. You can't point the feature at an existing table. Databricks creates a new one for you.
Checkpoint 3 of 5· Exam question
An ML engineer wants to build a new fine-tuning and evaluation dataset by combining real production traffic with SME-provided correct answers. They plan to join AI Gateway inference table records against a table of human-labeled ground truth. Which pair of inference table columns provides the raw content needed for this join?
Correct answer: A — `request` and `response`, which hold the raw JSON payloads sent to and returned from the endpoint
- A. The `request` and `response` columns store the actual JSON payloads exchanged with the endpoint, which is the content needed to pair a user's real question and the model's real answer with a ground-truth label for building a training or eval corpus.
- B. The identifier columns are useful for tracing and deduplication but contain no question or answer content, so they alone cannot supply the text needed to build a labeled dataset.
- C. HTTP status and requester identity describe the outcome and caller of a request, not the conversational content, so they don't provide the material needed to pair with SME labels.
- D. Timestamp and latency fields describe when and how fast a call ran, which supports performance analysis but does not contain the question or answer text required for a labeled dataset.
Checkpoint 4 of 5· Check yourself
You deploy a RAG agent programmatically, not with the agent deployment API, and the inference table has no traces. What is the documented fix?
Programmatic deployments switch on trace logging through an endpoint environment variable. Specifying an existing table is not supported at all.
“For programmatic deployments, set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.”Source: docs.databricks.com
Sources2
3.Reading the schema: which column answers which question
To assess performance from the log, you need to know what each column actually measures. The AI Gateway-enabled inference table for a serving endpoint has one row per request. Here are the columns you will use most:
| Column | What it holds | Use it to assess |
|---|---|---|
| databricks_request_id | Databricks-generated identifier attached to every request | Tracing one specific request end to end |
| request_date / request_time | UTC date and timestamp when the request was received | Trends over time, partitioning queries |
| status_code | HTTP status code returned from the model | Error rates |
| execution_duration_ms | Time the model spent on inference, excluding network overhead | Model compute time (not end-to-end latency) |
| request / response | Raw JSON bodies sent and returned | What users asked and what the app answered |
| sampling_fraction | Between 0 and 1; 1 means 100% of requests were included | Whether counts need scaling up |
| served_entity_id | Unique ID of the served entity | Joining to model details |
| requester | User or service principal whose permissions were used | Who is calling the app |
| logging_error_codes | Errors when data could not be logged, e.g. MAX_REQUEST_SIZE_EXCEEDED | Gaps in the log itself |
Two columns are easy to misread. execution_duration_ms covers only the model's own inference time. Network overhead isn't included, so it understates what the user actually waits. requester returns NULL for route-optimized custom model endpoints, so per-user analysis won't work on those.
Deployed agents get a payload request logs table built for conversations. As well as raw request and response strings, it pulls out fields that make RAG analysis much easier:
| Column | Description |
|---|---|
| conversation_id | The conversation id extracted from request logs |
| request | The last user query from the user's conversation |
| response | The last response to the user |
| request_raw / response_raw | String representation of the request / response |
| execution_time_ms | Time the model performed inference, excluding network overhead |
| trace | String representation of the trace |
Checkpoint 5 of 5· Match them up
Match each RAG performance question to the inference table column that answers it.
Tap a term, then the definition that fits it.
Each column records exactly one kind of measurement. execution_duration_ms in particular is model time only.
“This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”Source: docs.databricks.com
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.execution_duration_ms in a serving-endpoint inference table is the end-to-end latency the user experiences.Why is that wrong?
It measures only the model's inference time. Network overhead is excluded.
Covered in Reading the schema: which column answers which question
2.You can send inference logs to a Delta table you have already created, for example one that holds your evaluation set.Why is that wrong?
Databricks always creates a new inference table itself. Pointing at an existing table is not supported.
Covered in Turning on inference tables for an endpoint or agent
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-evaluation-monitoring-ragOfficial docs
“evaluation happens during development and monitoring happens once the application is deployed to production”
↩︎ Why a deployed RAG app needs its own logs“you need to know if the root cause of a low-quality answer is due to the retrieval step or a hallucination”
↩︎ Why a deployed RAG app needs its own logs“Your production logging must track the inputs, outputs, and intermediate steps such as document retrieval”
↩︎ Prediction“Once in production, you need to evaluate a significantly higher quantity of requests/responses and how each response was generated.”
↩︎ Checkpoint - 2.
“Provisioned throughput workloads Pay-per-token models External models Deployed agent Custom models”
↩︎ Turning on inference tables for an endpoint or agent“these inference tables store payload and request details as well as MLflow Trace logs.”
↩︎ Turning on inference tables for an endpoint or agent“Agents deployed using the mlflow.deploy() API have inference tables automatically enabled.”
↩︎ Turning on inference tables for an endpoint or agent“This field returns NULL for route-optimized custom model endpoints.”
↩︎ Reading the schema: which column answers which question“The last user query from the user's conversation.”
↩︎ Reading the schema: which column answers which question“The inference table automatically captures incoming requests and outgoing responses for an endpoint and logs them as a Unity Catalog Delta table.”
↩︎ Key concept“This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”
↩︎ Exam trap 1“Specifying an existing table is not supported.”
↩︎ Exam trap 2“In the AI Gateway section, select Enable inference tables.”
↩︎ Checkpoint“For programmatic deployments, set the ENABLE_MLFLOW_TRACING environment variable to True in the endpoint configuration.”
↩︎ Checkpoint“This does not include overhead network latencies and only represents the time it took for the model to generate predictions.”
↩︎ Checkpoint