What you will be able to do
- Describe what MLflow production monitoring does and what it requires before it can run
- Start a scorer for monitoring with register() and start(), and choose sample rates and trace filters
- Know monitoring's operating limits: the 20-scorer cap, the job identity, processing delay, and the SQL warehouse needed for Unity Catalog traces
- Compare the evaluation and monitoring phases by timing, trigger, data, coverage, output, and purpose
1.What the monitoring phase does after deployment
In the Gen AI life cycle, evaluation scores an app version against curated data during development. Monitoring takes over once the app is deployed. The RAG cookbook explains why a separate phase is needed: in production you have to assess a much larger volume of requests and responses, and how each response was generated. Production logging has to capture inputs, outputs, and intermediate steps such as document retrieval. Without them you can't tell whether a poor answer came from retrieval or from a hallucination.
MLflow production monitoring runs scorers continuously, in the background, on a sample of live production traces and attaches the results as feedback, so quality problems show up automatically after deployment. You schedule scorers against an MLflow experiment, and the monitoring service evaluates a configurable sample of incoming traces. It supports built-in and custom scorers, including multi-turn judges that evaluate entire conversations. For those, the assessment is attached to the first trace in each session. The feature is in Beta, and workspace admins control access from the Previews page.
Monitoring needs a few things in place first: an MLflow experiment where traces are logged, a production app instrumented with MLflow Tracing, tested scorers that work with the app's trace format, a serverless budget policy if the workspace doesn't allow the default one, and a SQL warehouse ID if traces are stored in Unity Catalog. The scorer requirement points back to the evaluation phase: if you used the production app as the predict_fn in mlflow.genai.evaluate() during development, your scorers are likely already compatible. Human reviewers also feed monitoring. The LLMOps guidance says human feedback should be managed like other data and fed into monitoring.
Checkpoint 1 of 6· Check yourself
Production traces for an agent are stored in Unity Catalog. Scorers have been registered and started, but no monitoring results appear. Which prerequisite is most likely missing?
For traces stored in Unity Catalog, monitoring needs a SQL warehouse ID. Monitoring doesn't take a predict_fn or an evaluation dataset; it scores live traces.
“If your traces are stored in Unity Catalog, you must configure a SQL warehouse ID for monitoring to work.”Source: docs.databricks.com
2.Starting monitors: register, start, and sample
Every scorer type follows the same two steps: built-in judges, custom judges, code-based scorers, and multi-turn judges. First register the scorer with the experiment, then start it with a sampling configuration. Registering alone does nothing; monitoring begins only when you start the scorer.
Checkpoint 2 of 6· Fill the gap
Which method begins continuous monitoring for this registered judge?
from mlflow.genai.scorers import Safety, ScorerSamplingConfig
# Register and start a built-in judge
safety_judge = Safety().register(name="safety")
safety_judge = safety_judge. ? (sampling_config=ScorerSamplingConfig(sample_rate=0.7))register() attaches the scorer to the experiment. start() with a ScorerSamplingConfig begins monitoring.
Source: docs.databricks.comEvaluation scores every row of its dataset. Monitoring usually doesn't score every trace: sampling lets you trade coverage against compute cost. The documentation gives guidance on rates:
| Scorer situation | Recommended sample_rate |
|---|---|
| Critical scorers such as safety and security checks | 1.0 |
| Expensive scorers, such as complex LLM judges | 0.05-0.2 |
| Iterative improvement during development | 0.3-0.5 |
ScorerSamplingConfig also takes a filter_string that narrows which traces a scorer sees. It uses the same filter syntax as mlflow.search_traces(), and you can combine conditions, for example successful traces from the last 24 hours.
from mlflow.genai.scorers import Safety, ScorerSamplingConfig
# Only evaluate traces that completed successfully
safety_judge = Safety().register(name="safety")
safety_judge = safety_judge.start(
sampling_config=ScorerSamplingConfig(
sample_rate=1.0,
filter_string="attributes.status = 'OK'"
),
)Some operating limits are worth knowing for the exam. At most 20 scorers can be associated with an experiment for continuous monitoring at any one time. The monitoring job runs under the identity of the user who first registered a scorer on the experiment, so that user's permissions decide what the job can access. After scheduling scorers, allow 15–20 minutes for initial processing. Then check the Traces tab for assessments and the monitoring dashboards for quality trends.
Checkpoint 3 of 6· Check yourself
A data engineer registered the first monitoring scorer on an experiment. Later, a teammate with broader permissions registers three more. Whose permissions govern what the monitoring job can access?
There is one monitoring job per experiment, and it runs as the user who first registered a scorer on that experiment.
“The monitoring job runs under the identity of the user who first registered a scorer on the experiment.”Source: docs.databricks.com
Sources3
3.Evaluation versus monitoring, side by side
The objective asks you to compare the two phases. Evaluation scores an app version in development against curated data. Monitoring scores sampled live traffic after deployment. Both use scorers, and both attach results to traces as feedback. The documentation describes MLflow Evaluation as the link between offline testing and production monitoring. Everything around the shared scorer is different:
| Dimension | Evaluation phase | Monitoring phase |
|---|---|---|
| When | During development, before release | Once the application is deployed to production |
| Trigger | An on-demand mlflow.genai.evaluate() run (or Run scorer in the UI) | Scorers registered with .register() and scheduled with .start() |
| Data | Evaluation dataset, pre-computed outputs, or existing traces | A sample of incoming live production traces |
| Coverage | Every record in the dataset | Controlled by sample_rate and filter_string |
| Output | An EvaluationResult and an evaluation run | Feedback on sampled traces, trends in monitoring dashboards |
| Purpose | Compare versions and prevent regressions before release | Catch regressions and detect new issues in live traffic |
Checkpoint 4 of 6· Exam question
A GenAI agent has been serving live customer traffic for three months. The team wants an ongoing, low-overhead way to track response quality trends without evaluating every single request, and to get near-real-time visibility into safety violations. What should they configure?
Correct answer: A — Scheduled scorers sample production traces logged via MLflow Tracing, using high sampling for safety checks and lower sampling for costlier judges
- A. This is correct: production monitoring registers scorers that run on a sampled portion of traces captured through MLflow Tracing, with sampling rates tuned per scorer so cheap or critical checks (like safety) can run at high coverage while expensive judges run at lower coverage. This gives ongoing quality visibility without scoring every request.
- B. A quarterly manual run against a static pre-launch dataset is characteristic of offline evaluation, not continuous production monitoring, and would leave long gaps where live quality drift or safety issues go undetected. It does not provide the near-real-time visibility the team needs.
- C. Conversation simulation generates synthetic test dialogues for development-time evaluation; it does not observe or score real production traffic. It cannot substitute for analyzing actual live traces.
- D. A ground-truth-dependent correctness judge needs labeled expected answers, which are not available for arbitrary live customer requests. Requiring labels for every production request is impractical and is not how continuous monitoring is set up.
Checkpoint 5 of 6· Match them up
Match each artifact to the phase role it plays.
Tap a term, then the definition that fits it.
evaluate() and evaluation datasets belong to the development phase. Sampling configuration and dashboards belong to continuous production monitoring.
“Configurable sampling rates so you can control the tradeoff between coverage and computational cost.”Source: docs.databricks.com
The two phases also feed each other. Monitoring that catches a regression in production starts the next turn of the loop: failing traces become new evaluation examples, a scorer is written or refined, the fix is evaluated, and the scorer goes back into monitoring.
Checkpoint 6 of 6· Exam question
An engineer wants to evaluate whether a RAG agent's retrieved context actually contains the facts needed to produce a labeled expected answer, using a small internal benchmark of question/expected-answer pairs. Which type of judge is appropriate, and why?
Correct answer: A — A judge requiring ground truth, such as retrieval sufficiency, because it compares retrieved context against the labeled answer to check fact coverage
- A. A retrieval sufficiency judge is a ground-truth-dependent judge that checks whether retrieved context contains the facts necessary to support the labeled expected answer, which is exactly what the benchmark with expected answers enables. This directly matches the stated goal of checking fact coverage against labels.
- B. A relevance-to-query judge is reference-free and assesses topical alignment between retrieved passages and the question, but it does not check whether the context contains the specific facts needed for a labeled expected answer. It would not use the benchmark's expected answers at all.
- C. A safety judge is reference-free and checks for harmful or inappropriate content, which is unrelated to whether retrieved context supports a labeled expected answer. It does not measure fact coverage.
- D. Tool call correctness is a ground-truth-dependent judge, but it validates agent tool invocations against expected tool outcomes, not whether retrieved context contains needed facts. It targets a different part of the agent's behavior than retrieval quality.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Production monitoring needs its own monitoring-specific metrics, separate from the scorers used in evaluation.Why is that wrong?
Monitoring reuses the evaluation scorers, so quality is measured the same way in development and production.
2.Registering a scorer with the experiment starts production monitoring.Why is that wrong?
Registering only attaches the scorer. Monitoring begins when you call start() with a sampling configuration.
3.Like evaluation, monitoring scores every production trace.Why is that wrong?
Monitoring evaluates a configurable sample of incoming traces. Expensive LLM judges are typically sampled at 5–20%, and only critical safety checks run at 1.0.
Practise it for real
Turn on production monitoring for an instrumented agent with a built-in Safety scorer and confirm that assessments appear on live traces.
1.Confirm the agent logs traces to an MLflow experiment. If the traces are stored in Unity Catalog, configure a SQL warehouse ID for the experiment.
Why: Monitoring scores traces in the experiment, and Unity Catalog traces need a SQL warehouse.
You should see: Recent production traces are visible in the experiment's Traces tab.
2.Run: safety = Safety().register(name="safety")
Why: Registering attaches the scorer to the experiment but doesn't start scoring yet.
You should see: A registered scorer named safety exists on the experiment; no feedback is added yet.
3.Run: safety = safety.start(sampling_config=ScorerSamplingConfig(sample_rate=1.0))
Why: Safety is a critical check, so it is sampled at 100%. Starting the scorer is what begins monitoring.
You should see: The scorer is now active for continuous monitoring.
4.Wait 15–20 minutes, then open the Traces tab and the monitoring dashboards.
Why: Initial processing takes time after scorers are scheduled.
You should see: Safety feedback is attached to sampled traces, and the dashboards show quality trends.
Stuck? Get a nudge
If no feedback appears after 20 minutes on Unity Catalog traces, check the SQL warehouse configuration first.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Production monitoring continuously runs MLflow scorers on a sample of your live production traces and attaches the results as feedback”
↩︎ What the monitoring phase does after deployment“the same scorers you built during evaluation now watch production”
↩︎ Prediction - 2.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-evaluation-monitoring-ragOfficial docs
“track the inputs, outputs, and intermediate steps such as document retrieval to enable ongoing monitoring”
↩︎ What the monitoring phase does after deployment“evaluation happens during development and monitoring happens once the application is deployed to production, but the fundamental components are similar”
↩︎ Evaluation versus monitoring, side by side - 3.
“If you used your production app as the predict_fn in mlflow.genai.evaluate() during development, your scorers are likely already compatible.”
↩︎ What the monitoring phase does after deployment“At any given time, at most 20 scorers can be associated with an experiment for continuous quality monitoring.”
↩︎ Starting monitors: register, start, and sample“After scheduling scorers, allow 15-20 minutes for initial processing.”
↩︎ Starting monitors: register, start, and sample“Use the same scorers in development and production to ensure consistent evaluation.”
↩︎ Exam trap 1“Monitoring begins the moment you start a scorer.”
↩︎ Exam trap 2“For expensive scorers, such as complex LLM judges, use lower sample rates (0.05-0.2).”
↩︎ Exam trap 3“If your traces are stored in Unity Catalog, you must configure a SQL warehouse ID for monitoring to work.”
↩︎ Checkpoint“The monitoring job runs under the identity of the user who first registered a scorer on the experiment.”
↩︎ Checkpoint“Configurable sampling rates so you can control the tradeoff between coverage and computational cost.”
↩︎ Checkpoint - 4.
“Human feedback should be managed like other data, ideally incorporated into monitoring based on near real-time streaming.”
↩︎ What the monitoring phase does after deployment - 5.
“MLflow Evaluation connects offline testing with production monitoring.”
↩︎ Evaluation versus monitoring, side by side - 6.
“Monitor production. Run your scorers against live traffic so the issue doesn't recur — which surfaces the next one to fix.”
↩︎ Evaluation versus monitoring, side by side