CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 52/56

    Agent Monitoring: Running Scorers on Live Production Traces

    Use inference tables and Agent Monitoring to track a live LLM endpoint

    10 min read
    1.79% of exam
    2 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Describe how production monitoring scores a sample of live agent traces and attaches the results as feedback
    • Register and start built-in or custom scorers with ScorerSamplingConfig
    • Choose sample rates and filter strings that balance coverage against cost
    • Configure the SQL warehouse needed to monitor traces stored in Unity Catalog
    • Write custom scorers that the remote monitoring service can serialize

    1.From payload logs to automated quality scores

    A request/response log shows you what a live agent did, but it does not judge whether the answers were any good. Production monitoring in MLflow 3, which this lesson calls Agent Monitoring, fills that gap. It runs MLflow scorers on a sample of the traces your deployed agent produces and attaches each result to the trace as feedback. These are the same scorers you used during evaluation, so development and production are measured by the same rules. The feature is in Beta, and workspace admins control access to it from the Previews page.

    Before you set it up, you need four things. First, an MLflow experiment that is receiving traces; if you do not specify one, the active experiment is used. Second, a production app instrumented with MLflow Tracing. Third, scorers you have already tested against your app's trace format. If you used your production app as the predict_fn in mlflow.genai.evaluate(), your scorers are probably compatible already. Fourth, if your workspace does not allow the default serverless budget policy, a budget policy set on the experiment before you register any scorers.

    Checkpoint 1 of 6· Check yourself

    What does production monitoring do with the result each time a scorer evaluates a live trace?

    Sources12

    2.The register-then-start pattern

    Monitoring is controlled per scorer, in two steps. .register() attaches the scorer to the experiment. .start() switches on scheduled scoring with a sampling configuration. Monitoring begins the moment a scorer is started. The same pattern works for built-in judges, custom judges, code-based scorers and multi-turn judges.

    Registering the built-in Safety judge and starting it on 70% of tracespython
    from mlflow.genai.scorers import Safety, ScorerSamplingConfig
    
    # Register and start a built-in judge
    safety_judge = Safety().register(name="safety")
    safety_judge = safety_judge.start(sampling_config=ScorerSamplingConfig(sample_rate=0.7))

    Checkpoint 2 of 6· Fill the gap

    Which method begins scheduled scoring of production traces after the scorer is registered?

    from mlflow.genai.scorers import Safety, ScorerSamplingConfig
    
    # Register a built-in scorer against your experiment, then start monitoring
    safety = Safety().register(name="safety")
    safety = safety. ? (sampling_config=ScorerSamplingConfig(sample_rate=1.0))

    There are a few operating limits to remember. An experiment can have at most 20 scorers doing continuous monitoring at any one time. After scheduling, allow 15–20 minutes for initial processing, then open the experiment's Traces tab to see the assessments and use the monitoring dashboards to follow quality over time. Multi-turn judges evaluate whole conversations, and their assessments are attached to the first trace in each session.

    Checkpoint 3 of 6· Exam question

    In what format and storage layer does Databricks persist the request and response data captured by an inference table for a Model Serving endpoint?

    Sources2

    3.Sampling rates and filter strings

    Every scored trace costs compute, and LLM judges cost the most. sample_rate sets the fraction of traces a scorer evaluates. The documentation gives guidance tied to how important and how expensive each scorer is.

    Recommended sample rates by scorer type
    Scorer situationRecommended sample_rate
    Critical scorers such as safety and security checks1.0
    Expensive scorers such as complex LLM judges0.05–0.2
    Iterative improvement during development0.3–0.5

    Checkpoint 4 of 6· Match them up

    Match each scorer to the sample rate Databricks recommends for it.

    Tap a term, then the definition that fits it.

    Sampling chooses how many traces get scored. filter_string chooses which ones. It uses the same syntax as mlflow.search_traces(), and you can combine conditions with AND, for example to score only successful traces from the last 24 hours.

    Restricting a scorer to traces that completed successfullypython
    from mlflow.genai.scorers import Safety, ScorerSamplingConfig
    
    # Only evaluate traces that completed successfully
    safety_judge = Safety().register(name="safety")
    safety_judge = safety_judge.start(
        sampling_config=ScorerSamplingConfig(
            sample_rate=1.0,
            filter_string="attributes.status = 'OK'"
        ),
    )

    Sources2

    4.Monitoring traces stored in Unity Catalog

    When traces live in Unity Catalog, the monitoring job queries those tables through a SQL warehouse. You must therefore save a warehouse ID on the experiment before registering scorers; otherwise the jobs fail with an error saying the mlflow.monitoring.sqlWarehouseId tag is missing. The helper below writes that tag for you.

    Saving the monitoring SQL warehouse ID on the experimentpython
    from mlflow.tracing import set_databricks_monitoring_sql_warehouse_id
    
    # Set the SQL warehouse ID for monitoring
    set_databricks_monitoring_sql_warehouse_id(
        sql_warehouse_id="<SQL_WAREHOUSE_ID>",
        experiment_id="<EXPERIMENT_ID>"  # Optional, uses active experiment if not specified
    )

    You need CAN USE on the warehouse and CAN EDIT on the experiment. Permission on the monitoring job is granted automatically when you register the first scorer. The job then runs as the user who registered that first scorer, so that person's access determines what monitoring can read.

    Checkpoint 5 of 6· Check yourself

    Whose permissions determine what the production monitoring job can access?

    Sources2

    5.Custom scorers that survive serialization

    Custom scorers do not run in your session. The monitoring service serializes them and runs them remotely, which brings four constraints. You must define and register them from a Databricks notebook. All imports must sit inside the function body, because references to anything outside the function are not captured. Only @scorer-decorated functions are allowed; class-based Scorer subclasses cannot be serialized. And type hints that need an import, such as List from typing, cause serialization to fail. The scorer should also handle missing data gracefully and return a consistent type.

    A custom scorer written to survive serialization: the import is inside the functionpython
    # Include imports in the function definition
    @scorer
    def good_scorer(outputs):
        import json  # Inside function
        return len(json.dumps(outputs))

    Checkpoint 6 of 6· Check yourself

    Which custom scorer can be registered for production monitoring?

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Setting MLFLOW_TRACING_SQL_WAREHOUSE_ID in your notebook is enough for monitoring traces stored in Unity Catalog.Why is that wrong?

      The monitoring job runs separately and reads the warehouse ID from the experiment tag. Set it with set_databricks_monitoring_sql_warehouse_id().

      Covered in Monitoring traces stored in Unity Catalog

    2. 2.Any scorer that works in mlflow.genai.evaluate(), including a class-based Scorer subclass, can be scheduled for production monitoring.Why is that wrong?

      Production scorers are serialized for remote execution, so only self-contained @scorer functions are supported.

      Covered in Custom scorers that survive serialization

    3. 3.To save money, every production scorer, including safety checks, should run on a small sample.Why is that wrong?

      Critical scorers such as safety and security checks should score every trace. Lower rates are for expensive judges.

      Covered in Sampling rates and filter strings

    Practise it for real

    Start a built-in Safety judge on a deployed agent's production traces and see its feedback in the MLflow experiment

    1. 1.If your traces are stored in Unity Catalog, call set_databricks_monitoring_sql_warehouse_id() with your warehouse ID and experiment ID.

      Why: The monitoring job reads the warehouse ID from the mlflow.monitoring.sqlWarehouseId experiment tag, not from your notebook's environment.

      You should see: The experiment carries the mlflow.monitoring.sqlWarehouseId tag.

    2. 2.Run Safety().register(name="safety") against the experiment that receives your agent's traces.

      Why: Registering attaches the scorer to the experiment and grants you permission on the monitoring job.

      You should see: A registered scorer object is returned; no traces are scored yet.

    3. 3.Call .start(sampling_config=ScorerSamplingConfig(sample_rate=1.0)) on the registered scorer.

      Why: Monitoring begins only when the scorer is started, and safety is a critical scorer that should see every trace.

      You should see: The scorer is active on new production traces.

    4. 4.Wait 15–20 minutes, then open the experiment's Traces tab.

      Why: Initial processing takes time after scheduling.

      You should see: Safety assessments are attached as feedback to the sampled traces.

    Stuck? Get a nudge

    If nothing appears, check that traces are logged to the experiment rather than to individual runs, and that any filter_string matches real traces.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “It closes the loop — the same scorers you built during evaluation now watch production.”
      ↩︎ From payload logs to automated quality scores
      “Production monitoring continuously runs MLflow scorers on a sample of your live production traces and attaches the results as feedback”
      ↩︎ Checkpoint
    2. 2.
      “Instrumented production application: Your agent must log traces using MLflow Tracing.”
      ↩︎ From payload logs to automated quality scores
      “Serverless budget policy: If your workspace does not allow the default serverless budget policy, set a policy on the MLflow experiment before registering scorers.”
      ↩︎ From payload logs to automated quality scores
      “This same two-step pattern works for any scorer type: built-in judges, custom judges, code-based scorers, and multi-turn judges.”
      ↩︎ The register-then-start pattern
      “At any given time, at most 20 scorers can be associated with an experiment for continuous quality monitoring.”
      ↩︎ The register-then-start pattern
      “For multi-turn judges, assessments are attached to the first trace in each session.”
      ↩︎ The register-then-start pattern
      “Use the filter_string parameter in ScorerSamplingConfig to control which traces a scorer evaluates.”
      ↩︎ Sampling rates and filter strings
      “Check experiment: Ensure that traces are logged to the experiment, not to individual runs.”
      ↩︎ Sampling rates and filter strings
      “Configure a warehouse ID on the experiment before you register scorers”
      ↩︎ Monitoring traces stored in Unity Catalog
      “CAN EDIT on the MLflow experiment.”
      ↩︎ Monitoring traces stored in Unity Catalog
      “Custom @scorer functions must be defined and registered from a Databricks notebook.”
      ↩︎ Custom scorers that survive serialization
      “Type hints in the function signature that require import statements (for example, List from typing) cause serialization failures.”
      ↩︎ Custom scorers that survive serialization
      “The monitoring job runs separately and reads the warehouse ID from the experiment tag.”
      ↩︎ Exam trap 1
      “Class-based Scorer subclasses cannot be serialized for remote execution.”
      ↩︎ Exam trap 2
      “For critical scorers such as safety and security checks, use sample_rate=1.0.”
      ↩︎ Exam trap 3
      “For expensive scorers, such as complex LLM judges, use lower sample rates (0.05-0.2).”
      ↩︎ Checkpoint
      “Setting the MLFLOW_TRACING_SQL_WAREHOUSE_ID environment variable in your notebook or application is not a substitute.”
      ↩︎ Prediction
      “The monitoring job runs under the identity of the user who first registered a scorer on the experiment.”
      ↩︎ Checkpoint
      “Class-based Scorer subclasses cannot be serialized for remote execution.”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.