What you will be able to do
- Explain how evaluation and production monitoring share the same metrics and scorers
- Pick retrieval, response, cost and latency metrics for a RAG or agent deployment
- Tell which metrics need human-labeled ground truth and which can score live traffic without it
- Tell component metrics, compound metrics and overall metrics apart
Key concept
Monitoring as continuous evaluation — Monitoring a deployed LLM app means measuring the same quality, cost and latency dimensions you measured during development, but on live traffic. So choosing monitoring metrics means choosing the metrics that still work once no labeled answer exists for each request.
1.Same metrics, different phase
Databricks describes evaluation and monitoring as two phases of one discipline. Both check whether a RAG application meets the quality, cost and latency requirements of its use case. Evaluation happens during development, against a curated evaluation set. Monitoring happens after deployment, against real requests. The Databricks guidance says the fundamental components of the two are similar.
In MLflow this continuity is concrete. A scorer is the unit that turns a trace into a quality assessment. It can be pass/fail, true/false, a number or a category, and the same scorer object can run in both phases. A scorer gets a trace from either evaluate() or the monitoring service. It extracts the fields it needs, assesses them, and returns the result as Feedback attached to that trace.
The phases differ in volume and in what you can see. In production you have far more requests than any evaluation set, and you need to diagnose individual bad answers. The cookbook therefore requires production trace logging that captures inputs, outputs and intermediate steps, not just the final answer. Without the retrieval step in the trace, you can't tell whether a poor answer came from the retriever or from the generator.
Checkpoint 1 of 5· Check yourself
A deployed RAG assistant gives a wrong answer. What must production logging capture so you can tell whether retrieval or hallucination caused it?
Root-cause diagnosis needs the intermediate retrieval step in the trace, not just the final output.
“Your production logging must track the inputs, outputs, and intermediate steps such as document retrieval”Source: docs.databricks.com
2.The four dimensions: retrieval, response, cost, latency
The Databricks RAG cookbook organises metrics into dimensions. Retrieval quality asks whether the app fetched relevant supporting data, and is built on precision and recall. Response quality asks whether the answer is accurate, grounded in the retrieved context (did the LLM hallucinate?) and safe. System performance covers cost and latency, measured as token consumption and end-to-end latency.
The cookbook insists you collect both retrieval and response metrics. A failure in one can be hidden by the other, so watching only the final answer leaves you unable to fix it.
| Dimension | Metric | Question it answers | Measured by | Needs ground truth? |
|---|---|---|---|---|
| Retrieval | chunk_relevance/precision | What % of the retrieved chunks are relevant to the request? | LLM judge | No |
| Retrieval | document_recall | What % of the ground truth documents are represented in the retrieved chunks? | Deterministic | Yes |
| Retrieval | context_sufficiency | Are the retrieved chunks sufficient to produce the expected response? | LLM judge | Yes |
| Response | correctness | Did the agent generate a correct response? | LLM judge | Yes |
| Response | relevance_to_query | Is the response relevant to the request? | LLM judge | No |
| Response | groundedness | Is the response a hallucination or grounded in context? | LLM judge | No |
| Response | safety | Is there harmful content in the response? | LLM judge | No |
| Cost | total_token_count, total_input_token_count, total_output_token_count | What's the total count of tokens for LLM generations? | Deterministic | No |
| Latency | latency_seconds | What's the latency of executing the app? | Deterministic | No |
The table also shows two ways of measuring. Cost and latency are deterministic: you compute them directly from the app's outputs, with no judgement involved. Most quality metrics come from an LLM judge, a separate model that reads the request, the retrieved context and the response and grades them. In MLflow these judges come as built-in scorers such as Correctness, RetrievalGroundedness and Safety. Custom LLM judges handle domain-specific criteria, and code-based scorers handle deterministic business logic.
Checkpoint 2 of 5· Check yourself
A team monitors only the correctness of final answers on its RAG chatbot. Why does Databricks recommend also tracking retrieval metrics?
Retrieval quality and response quality can diverge. Measuring both is the only way to locate the failing component.
“A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”Source: docs.databricks.com
Sources3
3.The deciding question: does the metric need ground truth?
The 'Needs ground truth?' column is the most useful column for monitoring. Some judges compare the app's output against a human-labeled expected answer. Others assess the output only from what is in the trace: the request, the retrieved context and the response.
During evaluation you have labels, so every row is available. A live production request arrives with no expected answer or list of relevant documents attached. Reference-free metrics fit that situation best: groundedness, relevance_to_query, safety and chunk precision for quality, plus the deterministic token-count and latency metrics. Correctness, document_recall and context_sufficiency depend on labels that live traffic doesn't provide.
The precision/recall pair shows why. Precision asks what share of the retrieved chunks are relevant, and a judge can decide that one chunk at a time. Recall asks what share of all relevant documents were retrieved. You can only answer that if you already know the full set of relevant documents.
Checkpoint 3 of 5· Check yourself
A support chatbot is live and nobody labels its answers. Which response-quality metric can still detect hallucinations on this traffic?
Groundedness checks the response against the retrieved context in the trace and needs no human-labeled answer. The other three all require ground truth.
“Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”Source: docs.databricks.com
Checkpoint 4 of 5· Exam question
A team deploys a customer support agent to production and wants to flag every safety violation, but can only afford to run their more expensive groundedness LLM judge against a small fraction of live traffic due to compute cost. How should they configure sample rates for their production monitoring scorers?
Correct answer: A — Set the safety scorer's sample rate to 1.0 so every request is checked, and set the groundedness scorer's sample rate to a smaller fraction such as 0.1 to control judge compute cost.
- A. Safety checks are typically critical and cheap enough to run at a sample rate of 1.0, while expensive judges like groundedness are sampled at a lower rate (e.g. 0.05-0.2) to control the cost of running an LLM judge. This matches the recommended tradeoff between coverage and computational cost in production monitoring.
- B. This inverts the priority: it under-samples safety, the check that should run on every request to catch violations reliably, while over-sampling the more expensive groundedness judge, which unnecessarily increases compute cost without improving safety coverage.
- C. An equal 0.5 sample rate for both scorers neither guarantees full coverage for the safety-critical check nor meaningfully reduces the cost of the expensive groundedness judge, so it satisfies neither goal well.
- D. Rate limits control request throughput and API cost at the gateway level, not the compute cost of running LLM-judge scorers against sampled traces, so this does not address the team's cost constraint on the groundedness judge.
Sources3
4.Component, compound and overall metrics
A classical ML model is one component, so its overall metrics are its component metrics. A RAG application chains several components together, and a change to one can affect the others. The cookbook separates three scopes of metric:
- Component metrics measure one stage on its own, for example precision @ K or nDCG for the retriever, or toxicity for the generator. - Compound metrics measure how stages interact. Faithfulness is the example: it checks whether the generator stuck to what the retriever returned, so it needs the chain input, the chain output and the retriever's output. - Overall metrics measure only the system's end-to-end input and output, for example answer correctness and latency.
Choosing metrics for a deployment means covering all three scopes. Which metrics matter most depends on the use case. The cookbook lists response accuracy, latency, cost and ratings from key stakeholders as examples.
Checkpoint 5 of 5· Match them up
Match each metric scope to an example from the Databricks RAG guidance
Tap a term, then the definition that fits it.
Component metrics look at one stage, compound metrics at how stages interact, and overall metrics at only the system's end-to-end input and output.
“Faithfulness measures the generator's adherence to the knowledge from a retriever that requires the chain input, chain output, and output of the internal retriever.”Source: docs.databricks.com
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Correctness works on unlabeled production traffic just as well as groundedness, because both are LLM judges.Why is that wrong?
Being an LLM judge doesn't make a metric reference-free. Correctness compares output against a human-labeled expected answer. Groundedness needs only the trace.
Covered in The deciding question: does the metric need ground truth?
2.Retrieval recall can be measured on any request, because the retrieved chunks are in the trace.Why is that wrong?
Recall needs to know every relevant document for the query, which means ground truth. Precision doesn't.
Covered in The deciding question: does the metric need ground truth?
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“You can use the same scorer for evaluation in development and monitoring in production to keep evaluation consistent throughout the application lifecycle.”
↩︎ Same metrics, different phase - 2.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-evaluation-monitoring-ragOfficial docs
“you need to know if the root cause of a low-quality answer is due to the retrieval step or a hallucination.”
↩︎ Same metrics, different phase“Depending on the application, important metrics might include response accuracy, latency, cost, or ratings from key stakeholders.”
↩︎ Component, compound and overall metrics“Technically, evaluation happens during development and monitoring happens once the application is deployed to production”
↩︎ Key concept“Your production logging must track the inputs, outputs, and intermediate steps such as document retrieval”
↩︎ Checkpoint“Faithfulness measures the generator's adherence to the knowledge from a retriever that requires the chain input, chain output, and output of the internal retriever.”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-assess-performanceOfficial docs
“Overall latency and token consumption are examples of chain performance metrics.”
↩︎ The four dimensions: retrieval, response, cost, latency“Cost and latency metrics can be computed deterministically based on the application's outputs.”
↩︎ The four dimensions: retrieval, response, cost, latency“Computing precision does not require knowing all relevant items.”
↩︎ The deciding question: does the metric need ground truth?“Some LLM judges, such as answer correctness, compare the human-labeled ground truth vs. the app outputs.”
↩︎ Exam trap 1“Computing recall requires your ground-truth to contain all relevant items.”
↩︎ Exam trap 2“A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
↩︎ Checkpoint“Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
↩︎ Checkpoint