What you will be able to do
- Schedule MLflow scorers on a sample of production traces and choose sampling rates by how critical and how expensive each scorer is
- Identify the endpoint health metrics Model Serving provides
- Use the Unity Gateway usage table and dashboard for token, latency, error-rate and cost metrics
- Tell when to use inference tables rather than aggregate usage metrics
1.Quality metrics on live traffic: scheduled scorers
Quality metrics such as safety and groundedness come from MLflow production monitoring, which is in Beta. You register scorers against the MLflow experiment where your agent logs traces. The monitoring service then runs those scorers on a configurable sample of incoming traces and attaches each result as feedback on the trace. Built-in judges, custom judges, code-based scorers and multi-turn judges all use the same .register() then .start() pattern.
from mlflow.genai.scorers import Safety, ScorerSamplingConfig
# Register and start a built-in judge
safety_judge = Safety().register(name="safety")
safety_judge = safety_judge.start(sampling_config=ScorerSamplingConfig(sample_rate=0.7))Sampling is how you trade coverage against compute cost. The guidance sets rates by how critical and how expensive a scorer is: sample_rate=1.0 for critical scorers such as safety and security checks, 0.05–0.2 for expensive LLM judges, and 0.3–0.5 for iterative improvement during development. You can also narrow the population with filter_string, which uses the same syntax as mlflow.search_traces(), for example to score only traces that completed successfully. An experiment can have at most 20 scorers for continuous monitoring at any one time. Allow 15–20 minutes after scheduling for results to appear in the Traces tab and the monitoring dashboards.
Checkpoint 1 of 5· Fill the gap
Which ScorerSamplingConfig parameter limits this safety judge to successfully completed traces?
from mlflow.genai.scorers import Safety, ScorerSamplingConfig
# Only evaluate traces that completed successfully
safety_judge = Safety().register(name="safety")
safety_judge = safety_judge.start(
sampling_config=ScorerSamplingConfig(
sample_rate=1.0,
? ="attributes.status = 'OK'"
),
)filter_string controls which traces a scorer evaluates, using the same filter syntax as mlflow.search_traces().
Checkpoint 2 of 5· Exam question
An ML engineer notices that response times for a deployed RAG serving endpoint have degraded over the past week and needs historical execution-duration and HTTP status data to investigate, without adding any extra logging code. Which built-in feature should they query?
Correct answer: A — The endpoint's inference table, since Databricks automatically logs execution duration, HTTP status codes, and request/response payloads for every serving endpoint request.
- A. Inference tables are Delta tables that automatically capture request and response payloads, execution duration, HTTP status codes, and timestamps for a serving endpoint, which is exactly the data needed to diagnose a latency regression without adding logging code.
- B. Unity Catalog audit logs record governance and access events such as who queried what, not per-request execution duration or HTTP status codes for model-serving traffic, so they are the wrong source for this investigation.
- C. AI Gateway usage tables are built for tracking token consumption and cost attribution across users and endpoints, not for capturing per-request execution duration and HTTP status codes, which is the role of inference tables instead.
- D. The MLflow tracking server stores experiment runs, parameters, and metrics logged explicitly by training or evaluation code, not automatically captured per-request latency and status codes from a live serving endpoint.
Sources1
2.Infrastructure metrics: endpoint health
Quality scorers don't show you whether the serving infrastructure itself is healthy. Model Serving provides that through endpoint health metrics: latency, request rate, error rate, CPU usage and memory usage. These are shown by default in the Serving UI for the last 14 days and can also be streamed to observability tools in real time.
The same Model Serving monitoring tools also include logs, which serve a different purpose. Ephemeral service logs capture stdout and stderr for debugging during deployment. Build logs help diagnose environment and dependency problems. Logs answer *why did this break?* Health metrics answer *how is the endpoint performing over time?*
Checkpoint 3 of 5· Check yourself
Which of these does Model Serving's endpoint health metrics tool report?
Endpoint health metrics cover infrastructure performance. Quality scores come from scorers, payloads from inference tables, and build output from build logs.
“Provides insights into infrastructure metrics like latency, request rate, error rate, CPU usage, and memory usage.”Source: docs.databricks.com
Sources2
3.Token, latency and cost metrics: Unity Gateway usage tracking
For models governed through Unity Gateway, the usage system table system.ai_gateway.usage records per-request metrics such as token counts, latency, requester and request tags. Usage tracking is a billed Gateway feature, and by default you need both account and metastore admin roles to query the table. One detail to remember: Gateway endpoints log to system.ai_gateway.usage, while legacy model serving endpoints log to system.serving.endpoint_usage.
The built-in usage dashboard groups these metrics into tabs that match the questions you ask about a deployment.
| Tab | Metrics tracked | Use it to |
|---|---|---|
| Overview | Daily request volume, token usage trends, top users by token consumption, unique user counts | Get a quick snapshot of activity |
| Performance | Latency percentiles (P50, P90, P95, P99), time to first byte, error rates, HTTP status code distributions | Spot performance bottlenecks or reliability issues |
| Usage | Token usage patterns, request distributions, cache hit ratios by model service, workspace and requester | Break down consumption |
| Cost Observability | Cost by model service, target model, user, service tags and request tags, plus estimated external model cost | Attribute and track spend |
For cost, the dashboard's Cost Analysis view combines the usage table with system.billing.usage and system.billing.list_prices to estimate foundation model spend. Estimated USD spend on external models is recorded hourly in system.ai_gateway.external_model_spend. Request or service tags let you attribute cost to a business unit or project.
The Performance tab of the Unity Gateway usage dashboard, which tracks the P50, P90, P95 and P99 latency percentiles. You can also query the latency column in system.ai_gateway.usage directly.
Checkpoint 4 of 5· Exam question
A company deploys a foundation model to power a high-traffic production chatbot and needs guaranteed throughput in tokens per second regardless of traffic spikes, rather than variable per-token billing. Which deployment approach and metric should they use to plan and monitor capacity?
Correct answer: A — Provisioned throughput, which reserves dedicated capacity measured in tokens per second with configured minimum and maximum throughput targets.
- A. Provisioned throughput lets a team reserve dedicated model-serving capacity measured in tokens per second, with configurable minimum and maximum targets, giving predictable throughput regardless of traffic spikes instead of variable per-token costs.
- B. Pay-per-token serving bills per token consumed and does not reserve dedicated throughput capacity; it is the variable-cost alternative the company is trying to move away from, not a mechanism for guaranteeing tokens-per-second throughput.
- C. Rate limiting caps the number of requests or tokens a caller can send per minute to protect shared capacity; it does not reserve dedicated inference capacity or guarantee a tokens-per-second throughput target for the model.
- D. Autoscaling compute clusters adjusts general-purpose cluster size based on load, but it is not the mechanism Databricks uses to reserve dedicated foundation-model inference capacity measured in tokens per second.
Sources3
4.Payload-level evidence: inference tables
Aggregate metrics tell you that something changed. Inference tables let you look at the requests behind the change. AI Gateway-enabled inference tables automatically log online prediction requests and responses to Unity Catalog Delta tables. They are available for endpoints serving custom models, external models or provisioned throughput workloads. Databricks names three uses: monitoring and debugging model quality or responses, generating training datasets, and conducting compliance audits. The Gateway observability guidance adds checks that guardrails such as PII redaction are working, and audits of sensitive-data exposure.
The division of work is: health metrics and the usage table show rates and percentiles, scorers grade quality, and inference tables hold the raw request/response evidence.
Checkpoint 5 of 5· Check yourself
A compliance team asks you to confirm that PII redaction is actually applied to responses from a Gateway-governed model. Which signal should you use?
Checking a guardrail means inspecting the actual payloads, and inference tables log those payloads. The other options only report aggregate performance numbers.
“Enable inference tables on a model service to log request and response payloads to a Unity Catalog table.”Source: docs.databricks.com
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Every production scorer should run at sample_rate=1.0 so no quality problem is missed.Why is that wrong?
Full coverage is recommended for critical checks such as safety. Expensive LLM judges should run on a small sample because sampling rate trades coverage against compute cost.
Covered in Quality metrics on live traffic: scheduled scorers
2.All serving endpoints log usage to the same system table.Why is that wrong?
Unity Gateway endpoints and legacy model serving endpoints write to different usage tables.
Covered in Token, latency and cost metrics: Unity Gateway usage tracking
Practise it for real
Start production quality monitoring on an agent's MLflow experiment, with sampling rates chosen per scorer
1.Confirm your agent logs traces to an MLflow experiment with MLflow Tracing. If the traces are stored in Unity Catalog, configure a SQL warehouse ID.
Why: Production scorers evaluate traces, so there has to be traced traffic to score.
You should see: Traces from recent requests are visible in the experiment's Traces tab.
2.Register the built-in Safety judge and start it with ScorerSamplingConfig(sample_rate=1.0).
Why: Safety is a critical check, and the guidance gives critical scorers full coverage.
You should see: The scorer is registered on the experiment and shown as started.
3.Register a more expensive judge, such as a custom LLM judge, and start it with a sample_rate between 0.05 and 0.2.
Why: Expensive judges run on a small sample to control compute cost.
You should see: Two scorers are now associated with the experiment, well under the 20-scorer limit.
4.Wait 15–20 minutes, then open the Traces tab and the monitoring dashboards.
Why: Initial processing takes time after scorers are scheduled.
You should see: Feedback from the scorers is attached to sampled traces, and quality trends appear on the dashboards.
Stuck? Get a nudge
If no assessments appear, check that the experiment has a serverless budget policy if your workspace doesn't allow the default one.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“For critical scorers such as safety and security checks, use sample_rate=1.0.”
↩︎ Quality metrics on live traffic: scheduled scorers“At any given time, at most 20 scorers can be associated with an experiment for continuous quality monitoring.”
↩︎ Quality metrics on live traffic: scheduled scorers“This same two-step pattern works for any scorer type: built-in judges, custom judges, code-based scorers, and multi-turn judges.”
↩︎ Quality metrics on live traffic: scheduled scorers“Configurable sampling rates so you can control the tradeoff between coverage and computational cost.”
↩︎ Exam trap 1“For expensive scorers, such as complex LLM judges, use lower sample rates (0.05-0.2).”
↩︎ Prediction - 2.https://docs.databricks.com/aws/en/machine-learning/model-serving/monitor-diagnose-endpointsOfficial docs
“Available by default in the Serving UI for the last 14 days.”
↩︎ Infrastructure metrics: endpoint health“Use this tool for monitoring and debugging model quality or responses, generating training data sets, or conducting compliance audits.”
↩︎ Payload-level evidence: inference tables“Provides insights into infrastructure metrics like latency, request rate, error rate, CPU usage, and memory usage.”
↩︎ Checkpoint - 3.
“Tracks key performance metrics including latency percentiles (P50, P90, P95, P99), time to first byte, error rates, and HTTP status code distributions.”
↩︎ Token, latency and cost metrics: Unity Gateway usage tracking“Usage tracking is a billed Unity Gateway feature.”
↩︎ Token, latency and cost metrics: Unity Gateway usage tracking
Also cited
“Unity Gateway endpoints log to system.ai_gateway.usage, whereas model serving endpoints (legacy) log to system.serving.endpoint_usage.”
↩︎ Exam trap 2“Enable inference tables on a model service to log request and response payloads to a Unity Catalog table.”
↩︎ Checkpoint