CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 48/56

    LLM Monitoring Metrics: Quality, Retrieval, Cost and Latency

    Select key metrics to monitor for a specific LLM deployment scenario

    9 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain how evaluation and production monitoring share the same metrics and scorers
    • Pick retrieval, response, cost and latency metrics for a RAG or agent deployment
    • Tell which metrics need human-labeled ground truth and which can score live traffic without it
    • Tell component metrics, compound metrics and overall metrics apart

    Key concept

    Monitoring as continuous evaluation — Monitoring a deployed LLM app means measuring the same quality, cost and latency dimensions you measured during development, but on live traffic. So choosing monitoring metrics means choosing the metrics that still work once no labeled answer exists for each request.

    1.Same metrics, different phase

    Databricks describes evaluation and monitoring as two phases of one discipline. Both check whether a RAG application meets the quality, cost and latency requirements of its use case. Evaluation happens during development, against a curated evaluation set. Monitoring happens after deployment, against real requests. The Databricks guidance says the fundamental components of the two are similar.

    In MLflow this continuity is concrete. A scorer is the unit that turns a trace into a quality assessment. It can be pass/fail, true/false, a number or a category, and the same scorer object can run in both phases. A scorer gets a trace from either evaluate() or the monitoring service. It extracts the fields it needs, assesses them, and returns the result as Feedback attached to that trace.

    The phases differ in volume and in what you can see. In production you have far more requests than any evaluation set, and you need to diagnose individual bad answers. The cookbook therefore requires production trace logging that captures inputs, outputs and intermediate steps, not just the final answer. Without the retrieval step in the trace, you can't tell whether a poor answer came from the retriever or from the generator.

    Checkpoint 1 of 5· Check yourself

    A deployed RAG assistant gives a wrong answer. What must production logging capture so you can tell whether retrieval or hallucination caused it?

    Sources12

    2.The four dimensions: retrieval, response, cost, latency

    The Databricks RAG cookbook organises metrics into dimensions. Retrieval quality asks whether the app fetched relevant supporting data, and is built on precision and recall. Response quality asks whether the answer is accurate, grounded in the retrieved context (did the LLM hallucinate?) and safe. System performance covers cost and latency, measured as token consumption and end-to-end latency.

    The cookbook insists you collect both retrieval and response metrics. A failure in one can be hidden by the other, so watching only the final answer leaves you unable to fix it.

    Metrics Databricks recommends for a RAG application, with how each is measured and whether it needs ground truth
    DimensionMetricQuestion it answersMeasured byNeeds ground truth?
    Retrievalchunk_relevance/precisionWhat % of the retrieved chunks are relevant to the request?LLM judgeNo
    Retrievaldocument_recallWhat % of the ground truth documents are represented in the retrieved chunks?DeterministicYes
    Retrievalcontext_sufficiencyAre the retrieved chunks sufficient to produce the expected response?LLM judgeYes
    ResponsecorrectnessDid the agent generate a correct response?LLM judgeYes
    Responserelevance_to_queryIs the response relevant to the request?LLM judgeNo
    ResponsegroundednessIs the response a hallucination or grounded in context?LLM judgeNo
    ResponsesafetyIs there harmful content in the response?LLM judgeNo
    Costtotal_token_count, total_input_token_count, total_output_token_countWhat's the total count of tokens for LLM generations?DeterministicNo
    Latencylatency_secondsWhat's the latency of executing the app?DeterministicNo

    The table also shows two ways of measuring. Cost and latency are deterministic: you compute them directly from the app's outputs, with no judgement involved. Most quality metrics come from an LLM judge, a separate model that reads the request, the retrieved context and the response and grades them. In MLflow these judges come as built-in scorers such as Correctness, RetrievalGroundedness and Safety. Custom LLM judges handle domain-specific criteria, and code-based scorers handle deterministic business logic.

    Checkpoint 2 of 5· Check yourself

    A team monitors only the correctness of final answers on its RAG chatbot. Why does Databricks recommend also tracking retrieval metrics?

    Sources3

    3.The deciding question: does the metric need ground truth?

    The 'Needs ground truth?' column is the most useful column for monitoring. Some judges compare the app's output against a human-labeled expected answer. Others assess the output only from what is in the trace: the request, the retrieved context and the response.

    During evaluation you have labels, so every row is available. A live production request arrives with no expected answer or list of relevant documents attached. Reference-free metrics fit that situation best: groundedness, relevance_to_query, safety and chunk precision for quality, plus the deterministic token-count and latency metrics. Correctness, document_recall and context_sufficiency depend on labels that live traffic doesn't provide.

    The precision/recall pair shows why. Precision asks what share of the retrieved chunks are relevant, and a judge can decide that one chunk at a time. Recall asks what share of all relevant documents were retrieved. You can only answer that if you already know the full set of relevant documents.

    Checkpoint 3 of 5· Check yourself

    A support chatbot is live and nobody labels its answers. Which response-quality metric can still detect hallucinations on this traffic?

    Checkpoint 4 of 5· Exam question

    A team deploys a customer support agent to production and wants to flag every safety violation, but can only afford to run their more expensive groundedness LLM judge against a small fraction of live traffic due to compute cost. How should they configure sample rates for their production monitoring scorers?

    Sources3

    4.Component, compound and overall metrics

    A classical ML model is one component, so its overall metrics are its component metrics. A RAG application chains several components together, and a change to one can affect the others. The cookbook separates three scopes of metric:

    - Component metrics measure one stage on its own, for example precision @ K or nDCG for the retriever, or toxicity for the generator. - Compound metrics measure how stages interact. Faithfulness is the example: it checks whether the generator stuck to what the retriever returned, so it needs the chain input, the chain output and the retriever's output. - Overall metrics measure only the system's end-to-end input and output, for example answer correctness and latency.

    Choosing metrics for a deployment means covering all three scopes. Which metrics matter most depends on the use case. The cookbook lists response accuracy, latency, cost and ratings from key stakeholders as examples.

    Checkpoint 5 of 5· Match them up

    Match each metric scope to an example from the Databricks RAG guidance

    Tap a term, then the definition that fits it.

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Correctness works on unlabeled production traffic just as well as groundedness, because both are LLM judges.Why is that wrong?

      Being an LLM judge doesn't make a metric reference-free. Correctness compares output against a human-labeled expected answer. Groundedness needs only the trace.

      Covered in The deciding question: does the metric need ground truth?

    2. 2.Retrieval recall can be measured on any request, because the retrieved chunks are in the trace.Why is that wrong?

      Recall needs to know every relevant document for the query, which means ground truth. Precision doesn't.

      Covered in The deciding question: does the metric need ground truth?

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “You can use the same scorer for evaluation in development and monitoring in production to keep evaluation consistent throughout the application lifecycle.”
      ↩︎ Same metrics, different phase
    2. 2.
      “you need to know if the root cause of a low-quality answer is due to the retrieval step or a hallucination.”
      ↩︎ Same metrics, different phase
      “Depending on the application, important metrics might include response accuracy, latency, cost, or ratings from key stakeholders.”
      ↩︎ Component, compound and overall metrics
      “Technically, evaluation happens during development and monitoring happens once the application is deployed to production”
      ↩︎ Key concept
      “Your production logging must track the inputs, outputs, and intermediate steps such as document retrieval”
      ↩︎ Checkpoint
      “Faithfulness measures the generator's adherence to the knowledge from a retriever that requires the chain input, chain output, and output of the internal retriever.”
      ↩︎ Checkpoint
    3. 3.
      “Overall latency and token consumption are examples of chain performance metrics.”
      ↩︎ The four dimensions: retrieval, response, cost, latency
      “Cost and latency metrics can be computed deterministically based on the application's outputs.”
      ↩︎ The four dimensions: retrieval, response, cost, latency
      “Computing precision does not require knowing all relevant items.”
      ↩︎ The deciding question: does the metric need ground truth?
      “Some LLM judges, such as answer correctness, compare the human-labeled ground truth vs. the app outputs.”
      ↩︎ Exam trap 1
      “Computing recall requires your ground-truth to contain all relevant items.”
      ↩︎ Exam trap 2
      “A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
      ↩︎ Checkpoint
      “Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Production LLM Monitoring Signals on Databricks

    Spotted a mistake, or was something unclear? Tell us.