CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 26/56

    GenAI Production Monitoring vs Evaluation in MLflow

    Compare the evaluation and monitoring phases of the Gen AI application life cycle

    10 min read
    1.79% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Describe what MLflow production monitoring does and what it requires before it can run
    • Start a scorer for monitoring with register() and start(), and choose sample rates and trace filters
    • Know monitoring's operating limits: the 20-scorer cap, the job identity, processing delay, and the SQL warehouse needed for Unity Catalog traces
    • Compare the evaluation and monitoring phases by timing, trigger, data, coverage, output, and purpose

    1.What the monitoring phase does after deployment

    In the Gen AI life cycle, evaluation scores an app version against curated data during development. Monitoring takes over once the app is deployed. The RAG cookbook explains why a separate phase is needed: in production you have to assess a much larger volume of requests and responses, and how each response was generated. Production logging has to capture inputs, outputs, and intermediate steps such as document retrieval. Without them you can't tell whether a poor answer came from retrieval or from a hallucination.

    MLflow production monitoring runs scorers continuously, in the background, on a sample of live production traces and attaches the results as feedback, so quality problems show up automatically after deployment. You schedule scorers against an MLflow experiment, and the monitoring service evaluates a configurable sample of incoming traces. It supports built-in and custom scorers, including multi-turn judges that evaluate entire conversations. For those, the assessment is attached to the first trace in each session. The feature is in Beta, and workspace admins control access from the Previews page.

    Monitoring needs a few things in place first: an MLflow experiment where traces are logged, a production app instrumented with MLflow Tracing, tested scorers that work with the app's trace format, a serverless budget policy if the workspace doesn't allow the default one, and a SQL warehouse ID if traces are stored in Unity Catalog. The scorer requirement points back to the evaluation phase: if you used the production app as the predict_fn in mlflow.genai.evaluate() during development, your scorers are likely already compatible. Human reviewers also feed monitoring. The LLMOps guidance says human feedback should be managed like other data and fed into monitoring.

    Checkpoint 1 of 6· Check yourself

    Production traces for an agent are stored in Unity Catalog. Scorers have been registered and started, but no monitoring results appear. Which prerequisite is most likely missing?

    Sources1234

    2.Starting monitors: register, start, and sample

    Every scorer type follows the same two steps: built-in judges, custom judges, code-based scorers, and multi-turn judges. First register the scorer with the experiment, then start it with a sampling configuration. Registering alone does nothing; monitoring begins only when you start the scorer.

    Checkpoint 2 of 6· Fill the gap

    Which method begins continuous monitoring for this registered judge?

    from mlflow.genai.scorers import Safety, ScorerSamplingConfig
    
    # Register and start a built-in judge
    safety_judge = Safety().register(name="safety")
    safety_judge = safety_judge. ? (sampling_config=ScorerSamplingConfig(sample_rate=0.7))

    Evaluation scores every row of its dataset. Monitoring usually doesn't score every trace: sampling lets you trade coverage against compute cost. The documentation gives guidance on rates:

    Recommended sample_rate by scorer type in production monitoring
    Scorer situationRecommended sample_rate
    Critical scorers such as safety and security checks1.0
    Expensive scorers, such as complex LLM judges0.05-0.2
    Iterative improvement during development0.3-0.5

    ScorerSamplingConfig also takes a filter_string that narrows which traces a scorer sees. It uses the same filter syntax as mlflow.search_traces(), and you can combine conditions, for example successful traces from the last 24 hours.

    Restricting a monitoring scorer to successfully completed traces with filter_stringpython
    from mlflow.genai.scorers import Safety, ScorerSamplingConfig
    
    # Only evaluate traces that completed successfully
    safety_judge = Safety().register(name="safety")
    safety_judge = safety_judge.start(
        sampling_config=ScorerSamplingConfig(
            sample_rate=1.0,
            filter_string="attributes.status = 'OK'"
        ),
    )

    Some operating limits are worth knowing for the exam. At most 20 scorers can be associated with an experiment for continuous monitoring at any one time. The monitoring job runs under the identity of the user who first registered a scorer on the experiment, so that user's permissions decide what the job can access. After scheduling scorers, allow 15–20 minutes for initial processing. Then check the Traces tab for assessments and the monitoring dashboards for quality trends.

    Checkpoint 3 of 6· Check yourself

    A data engineer registered the first monitoring scorer on an experiment. Later, a teammate with broader permissions registers three more. Whose permissions govern what the monitoring job can access?

    Sources3

    3.Evaluation versus monitoring, side by side

    The objective asks you to compare the two phases. Evaluation scores an app version in development against curated data. Monitoring scores sampled live traffic after deployment. Both use scorers, and both attach results to traces as feedback. The documentation describes MLflow Evaluation as the link between offline testing and production monitoring. Everything around the shared scorer is different:

    Evaluation phase vs monitoring phase in the MLflow GenAI life cycle
    DimensionEvaluation phaseMonitoring phase
    WhenDuring development, before releaseOnce the application is deployed to production
    TriggerAn on-demand mlflow.genai.evaluate() run (or Run scorer in the UI)Scorers registered with .register() and scheduled with .start()
    DataEvaluation dataset, pre-computed outputs, or existing tracesA sample of incoming live production traces
    CoverageEvery record in the datasetControlled by sample_rate and filter_string
    OutputAn EvaluationResult and an evaluation runFeedback on sampled traces, trends in monitoring dashboards
    PurposeCompare versions and prevent regressions before releaseCatch regressions and detect new issues in live traffic

    Checkpoint 4 of 6· Exam question

    A GenAI agent has been serving live customer traffic for three months. The team wants an ongoing, low-overhead way to track response quality trends without evaluating every single request, and to get near-real-time visibility into safety violations. What should they configure?

    Checkpoint 5 of 6· Match them up

    Match each artifact to the phase role it plays.

    Tap a term, then the definition that fits it.

    The two phases also feed each other. Monitoring that catches a regression in production starts the next turn of the loop: failing traces become new evaluation examples, a scorer is written or refined, the fix is evaluated, and the scorer goes back into monitoring.

    Checkpoint 6 of 6· Exam question

    An engineer wants to evaluate whether a RAG agent's retrieved context actually contains the facts needed to produce a labeled expected answer, using a small internal benchmark of question/expected-answer pairs. Which type of judge is appropriate, and why?

    Sources256

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Production monitoring needs its own monitoring-specific metrics, separate from the scorers used in evaluation.Why is that wrong?

      Monitoring reuses the evaluation scorers, so quality is measured the same way in development and production.

      Covered in What the monitoring phase does after deployment

    2. 2.Registering a scorer with the experiment starts production monitoring.Why is that wrong?

      Registering only attaches the scorer. Monitoring begins when you call start() with a sampling configuration.

      Covered in Starting monitors: register, start, and sample

    3. 3.Like evaluation, monitoring scores every production trace.Why is that wrong?

      Monitoring evaluates a configurable sample of incoming traces. Expensive LLM judges are typically sampled at 5–20%, and only critical safety checks run at 1.0.

      Covered in Starting monitors: register, start, and sample

    Practise it for real

    Turn on production monitoring for an instrumented agent with a built-in Safety scorer and confirm that assessments appear on live traces.

    1. 1.Confirm the agent logs traces to an MLflow experiment. If the traces are stored in Unity Catalog, configure a SQL warehouse ID for the experiment.

      Why: Monitoring scores traces in the experiment, and Unity Catalog traces need a SQL warehouse.

      You should see: Recent production traces are visible in the experiment's Traces tab.

    2. 2.Run: safety = Safety().register(name="safety")

      Why: Registering attaches the scorer to the experiment but doesn't start scoring yet.

      You should see: A registered scorer named safety exists on the experiment; no feedback is added yet.

    3. 3.Run: safety = safety.start(sampling_config=ScorerSamplingConfig(sample_rate=1.0))

      Why: Safety is a critical check, so it is sampled at 100%. Starting the scorer is what begins monitoring.

      You should see: The scorer is now active for continuous monitoring.

    4. 4.Wait 15–20 minutes, then open the Traces tab and the monitoring dashboards.

      Why: Initial processing takes time after scorers are scheduled.

      You should see: Safety feedback is attached to sampled traces, and the dashboards show quality trends.

    Stuck? Get a nudge

    If no feedback appears after 20 minutes on Unity Catalog traces, check the SQL warehouse configuration first.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Production monitoring continuously runs MLflow scorers on a sample of your live production traces and attaches the results as feedback”
      ↩︎ What the monitoring phase does after deployment
      “the same scorers you built during evaluation now watch production”
      ↩︎ Prediction
    2. 2.
      “track the inputs, outputs, and intermediate steps such as document retrieval to enable ongoing monitoring”
      ↩︎ What the monitoring phase does after deployment
      “evaluation happens during development and monitoring happens once the application is deployed to production, but the fundamental components are similar”
      ↩︎ Evaluation versus monitoring, side by side
    3. 3.
      “If you used your production app as the predict_fn in mlflow.genai.evaluate() during development, your scorers are likely already compatible.”
      ↩︎ What the monitoring phase does after deployment
      “At any given time, at most 20 scorers can be associated with an experiment for continuous quality monitoring.”
      ↩︎ Starting monitors: register, start, and sample
      “After scheduling scorers, allow 15-20 minutes for initial processing.”
      ↩︎ Starting monitors: register, start, and sample
      “Use the same scorers in development and production to ensure consistent evaluation.”
      ↩︎ Exam trap 1
      “Monitoring begins the moment you start a scorer.”
      ↩︎ Exam trap 2
      “For expensive scorers, such as complex LLM judges, use lower sample rates (0.05-0.2).”
      ↩︎ Exam trap 3
      “If your traces are stored in Unity Catalog, you must configure a SQL warehouse ID for monitoring to work.”
      ↩︎ Checkpoint
      “The monitoring job runs under the identity of the user who first registered a scorer on the experiment.”
      ↩︎ Checkpoint
      “Configurable sampling rates so you can control the tradeoff between coverage and computational cost.”
      ↩︎ Checkpoint
    4. 4.
      “Human feedback should be managed like other data, ideally incorporated into monitoring based on near real-time streaming.”
      ↩︎ What the monitoring phase does after deployment
    5. 5.
      “MLflow Evaluation connects offline testing with production monitoring.”
      ↩︎ Evaluation versus monitoring, side by side
    6. 6.
      “Monitor production. Run your scorers against live traffic so the issue doesn't recur — which surfaces the next one to fix.”
      ↩︎ Evaluation versus monitoring, side by side

    Spotted a mistake, or was something unclear? Tell us.