CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 48/56

    Production LLM Monitoring Signals on Databricks

    Select key metrics to monitor for a specific LLM deployment scenario

    9 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Schedule MLflow scorers on a sample of production traces and choose sampling rates by how critical and how expensive each scorer is
    • Identify the endpoint health metrics Model Serving provides
    • Use the Unity Gateway usage table and dashboard for token, latency, error-rate and cost metrics
    • Tell when to use inference tables rather than aggregate usage metrics

    1.Quality metrics on live traffic: scheduled scorers

    Quality metrics such as safety and groundedness come from MLflow production monitoring, which is in Beta. You register scorers against the MLflow experiment where your agent logs traces. The monitoring service then runs those scorers on a configurable sample of incoming traces and attaches each result as feedback on the trace. Built-in judges, custom judges, code-based scorers and multi-turn judges all use the same .register() then .start() pattern.

    Register a built-in Safety judge and start it on 70% of production tracespython
    from mlflow.genai.scorers import Safety, ScorerSamplingConfig
    
    # Register and start a built-in judge
    safety_judge = Safety().register(name="safety")
    safety_judge = safety_judge.start(sampling_config=ScorerSamplingConfig(sample_rate=0.7))

    Sampling is how you trade coverage against compute cost. The guidance sets rates by how critical and how expensive a scorer is: sample_rate=1.0 for critical scorers such as safety and security checks, 0.05–0.2 for expensive LLM judges, and 0.3–0.5 for iterative improvement during development. You can also narrow the population with filter_string, which uses the same syntax as mlflow.search_traces(), for example to score only traces that completed successfully. An experiment can have at most 20 scorers for continuous monitoring at any one time. Allow 15–20 minutes after scheduling for results to appear in the Traces tab and the monitoring dashboards.

    Checkpoint 1 of 5· Fill the gap

    Which ScorerSamplingConfig parameter limits this safety judge to successfully completed traces?

    from mlflow.genai.scorers import Safety, ScorerSamplingConfig
    
    # Only evaluate traces that completed successfully
    safety_judge = Safety().register(name="safety")
    safety_judge = safety_judge.start(
        sampling_config=ScorerSamplingConfig(
            sample_rate=1.0,
             ? ="attributes.status = 'OK'"
        ),
    )

    Checkpoint 2 of 5· Exam question

    An ML engineer notices that response times for a deployed RAG serving endpoint have degraded over the past week and needs historical execution-duration and HTTP status data to investigate, without adding any extra logging code. Which built-in feature should they query?

    Sources1

    2.Infrastructure metrics: endpoint health

    Quality scorers don't show you whether the serving infrastructure itself is healthy. Model Serving provides that through endpoint health metrics: latency, request rate, error rate, CPU usage and memory usage. These are shown by default in the Serving UI for the last 14 days and can also be streamed to observability tools in real time.

    The same Model Serving monitoring tools also include logs, which serve a different purpose. Ephemeral service logs capture stdout and stderr for debugging during deployment. Build logs help diagnose environment and dependency problems. Logs answer *why did this break?* Health metrics answer *how is the endpoint performing over time?*

    Checkpoint 3 of 5· Check yourself

    Which of these does Model Serving's endpoint health metrics tool report?

    Sources2

    3.Token, latency and cost metrics: Unity Gateway usage tracking

    For models governed through Unity Gateway, the usage system table system.ai_gateway.usage records per-request metrics such as token counts, latency, requester and request tags. Usage tracking is a billed Gateway feature, and by default you need both account and metastore admin roles to query the table. One detail to remember: Gateway endpoints log to system.ai_gateway.usage, while legacy model serving endpoints log to system.serving.endpoint_usage.

    The built-in usage dashboard groups these metrics into tabs that match the questions you ask about a deployment.

    Unity Gateway usage dashboard tabs and the metrics each one tracks
    TabMetrics trackedUse it to
    OverviewDaily request volume, token usage trends, top users by token consumption, unique user countsGet a quick snapshot of activity
    PerformanceLatency percentiles (P50, P90, P95, P99), time to first byte, error rates, HTTP status code distributionsSpot performance bottlenecks or reliability issues
    UsageToken usage patterns, request distributions, cache hit ratios by model service, workspace and requesterBreak down consumption
    Cost ObservabilityCost by model service, target model, user, service tags and request tags, plus estimated external model costAttribute and track spend

    For cost, the dashboard's Cost Analysis view combines the usage table with system.billing.usage and system.billing.list_prices to estimate foundation model spend. Estimated USD spend on external models is recorded hourly in system.ai_gateway.external_model_spend. Request or service tags let you attribute cost to a business unit or project.

    Checkpoint 4 of 5· Exam question

    A company deploys a foundation model to power a high-traffic production chatbot and needs guaranteed throughput in tokens per second regardless of traffic spikes, rather than variable per-token billing. Which deployment approach and metric should they use to plan and monitor capacity?

    Sources3

    4.Payload-level evidence: inference tables

    Aggregate metrics tell you that something changed. Inference tables let you look at the requests behind the change. AI Gateway-enabled inference tables automatically log online prediction requests and responses to Unity Catalog Delta tables. They are available for endpoints serving custom models, external models or provisioned throughput workloads. Databricks names three uses: monitoring and debugging model quality or responses, generating training datasets, and conducting compliance audits. The Gateway observability guidance adds checks that guardrails such as PII redaction are working, and audits of sensitive-data exposure.

    The division of work is: health metrics and the usage table show rates and percentiles, scorers grade quality, and inference tables hold the raw request/response evidence.

    Checkpoint 5 of 5· Check yourself

    A compliance team asks you to confirm that PII redaction is actually applied to responses from a Gateway-governed model. Which signal should you use?

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Every production scorer should run at sample_rate=1.0 so no quality problem is missed.Why is that wrong?

      Full coverage is recommended for critical checks such as safety. Expensive LLM judges should run on a small sample because sampling rate trades coverage against compute cost.

      Covered in Quality metrics on live traffic: scheduled scorers

    2. 2.All serving endpoints log usage to the same system table.Why is that wrong?

      Unity Gateway endpoints and legacy model serving endpoints write to different usage tables.

      Covered in Token, latency and cost metrics: Unity Gateway usage tracking

    Practise it for real

    Start production quality monitoring on an agent's MLflow experiment, with sampling rates chosen per scorer

    1. 1.Confirm your agent logs traces to an MLflow experiment with MLflow Tracing. If the traces are stored in Unity Catalog, configure a SQL warehouse ID.

      Why: Production scorers evaluate traces, so there has to be traced traffic to score.

      You should see: Traces from recent requests are visible in the experiment's Traces tab.

    2. 2.Register the built-in Safety judge and start it with ScorerSamplingConfig(sample_rate=1.0).

      Why: Safety is a critical check, and the guidance gives critical scorers full coverage.

      You should see: The scorer is registered on the experiment and shown as started.

    3. 3.Register a more expensive judge, such as a custom LLM judge, and start it with a sample_rate between 0.05 and 0.2.

      Why: Expensive judges run on a small sample to control compute cost.

      You should see: Two scorers are now associated with the experiment, well under the 20-scorer limit.

    4. 4.Wait 15–20 minutes, then open the Traces tab and the monitoring dashboards.

      Why: Initial processing takes time after scorers are scheduled.

      You should see: Feedback from the scorers is attached to sampled traces, and quality trends appear on the dashboards.

    Stuck? Get a nudge

    If no assessments appear, check that the experiment has a serverless budget policy if your workspace doesn't allow the default one.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “For critical scorers such as safety and security checks, use sample_rate=1.0.”
      ↩︎ Quality metrics on live traffic: scheduled scorers
      “At any given time, at most 20 scorers can be associated with an experiment for continuous quality monitoring.”
      ↩︎ Quality metrics on live traffic: scheduled scorers
      “This same two-step pattern works for any scorer type: built-in judges, custom judges, code-based scorers, and multi-turn judges.”
      ↩︎ Quality metrics on live traffic: scheduled scorers
      “Configurable sampling rates so you can control the tradeoff between coverage and computational cost.”
      ↩︎ Exam trap 1
      “For expensive scorers, such as complex LLM judges, use lower sample rates (0.05-0.2).”
      ↩︎ Prediction
    2. 2.
      “Available by default in the Serving UI for the last 14 days.”
      ↩︎ Infrastructure metrics: endpoint health
      “Use this tool for monitoring and debugging model quality or responses, generating training data sets, or conducting compliance audits.”
      ↩︎ Payload-level evidence: inference tables
      “Provides insights into infrastructure metrics like latency, request rate, error rate, CPU usage, and memory usage.”
      ↩︎ Checkpoint
    3. 3.
      “Tracks key performance metrics including latency percentiles (P50, P90, P95, P99), time to first byte, error rates, and HTTP status code distributions.”
      ↩︎ Token, latency and cost metrics: Unity Gateway usage tracking
      “Usage tracking is a billed Unity Gateway feature.”
      ↩︎ Token, latency and cost metrics: Unity Gateway usage tracking

    Also cited

    Spotted a mistake, or was something unclear? Tell us.