What you will be able to do
- Iterate on a scorer against stored traces without rerunning the app
- Write scorers that inspect trace spans or wrap a built-in LLM judge
- Handle missing expectations and scorer errors correctly
- Know which scorer definitions production monitoring supports, and register one that reads secrets
1.Iterating on a scorer against stored traces
A code-based scorer is a Python function decorated with @scorer that receives an MLflow trace and returns a value or Feedback. To get it right you usually need several attempts, and calling an LLM-backed app on every attempt is slow. Databricks recommends a workflow that runs the app once and reuses its traces:
1. Define evaluation data. 2. Generate traces from your app. 3. Query and store the resulting traces. 4. Iterate on the scorer against the stored traces.
After that run, mlflow.search_traces(run_id=eval_results.run_id) returns the traces as a Pandas DataFrame. You pass that DataFrame straight back to evaluate() as data and leave out predict_fn. Only the new scorer runs, and the app is not called.
from mlflow.genai.scorers import scorer
@scorer
def response_length(outputs: str) -> int:
# Example metric.
# Implement your actual metric logic here.
return len(outputs)
# Note the lack of a predict_fn parameter.
mlflow.genai.evaluate(
data=generated_traces,
scorers=[response_length],
)Checkpoint 1 of 8· Put it in order
Put the scorer development workflow in order
- 1.Run evaluate() on the stored traces with the scorer you are iterating on
- 2.Generate traces by running evaluate() with predict_fn and a placeholder scorer
- 3.Define evaluation data
- 4.Store the traces with mlflow.search_traces()
You generate traces once, then evaluate the stored DataFrame repeatedly, so each change to the scorer avoids a new app run.
“This allows you to quickly iterate on your metric without having to re-run your app.”Source: docs.databricks.com
Sources1
2.Two common patterns: reading spans and wrapping a judge
When the scorer declares the trace argument, it can read individual spans with their inputs, outputs, attributes and timing. The documented example finds the chat-model span and checks that the LLM call finished within 5 seconds. It returns "yes" or "no" as Feedback, with a rationale that states the measured time.
Checkpoint 2 of 8· Fill the gap
Which Trace method finds the chat-model span in this latency scorer?
@scorer
def llm_response_time_good(trace: Trace) -> Feedback:
# Search particular span type from the trace
llm_span = trace. ? (span_type=SpanType.CHAT_MODEL)[0]
response_time = (llm_span.end_time_ns - llm_span.start_time_ns) / 1e9 # convert to secondstrace.search_spans(span_type=...) returns the matching spans. The scorer takes the first one and computes its duration from end_time_ns and start_time_ns.
Source: docs.databricks.comThe second pattern wraps a built-in LLM judge. Use it to preprocess trace data for the judge or to post-process the judge's feedback. In the example, the app's inputs is a whole chat history: {"messages": [...]}. The relevance judge needs a single request, so the scorer walks the messages backwards to find the last user message. It then calls is_context_relevant and returns that judge's Feedback unchanged.
if not last_user_message_content:
raise Exception("Could not extract the last user message from inputs to evaluate relevance.")
# Call the `relevance_to_query judge. It will return a Feedback object.
return is_context_relevant(
request=last_user_message_content,
context={"response": outputs},
)Sources2
3.Expectations: available offline, missing in production
Expectations are ground-truth values or labels. When you run mlflow.genai.evaluate(), you can supply them in two ways. One is an expectations key on each row of a list or DataFrame. The other is a trace field, as in the DataFrame from mlflow.search_traces(), which carries any Expectation data attached to the traces.
Production monitoring works differently. Registered scorers always parse inputs and outputs from the trace, and expectations is not available, because live traffic has no ground truth. A scorer that you reuse in both places must not assume that expectations is filled in.
Checkpoint 3 of 8· Check yourself
A scorer compares outputs with expectations["expected_response"]. It passes offline evaluation, and you then register it for production monitoring. What should you expect?
Registered production scorers parse only inputs and outputs from the trace. Design the scorer to handle missing expectations gracefully.
“Registered scorers for production monitoring always parse the inputs and outputs parameters from the trace. expectations is not available.”Source: docs.databricks.com
Checkpoint 4 of 8· Exam question
A team writes a custom scorer for a support agent and wants stakeholders reviewing evaluation runs to see not just a numeric score but also a short written reason for that score, attached to each trace. What should the `@scorer`-decorated function return to achieve this?
Correct answer: A — Return a `Feedback` object carrying a numeric `value` plus a `rationale` string explaining why that score was assigned
- A. A `Feedback` object is built for exactly this case: it pairs a `value` (the score) with a `rationale` (the explanation), giving stakeholders both the number and the reasoning behind it in a structured way attached to the trace.
- B. A plain `float` is a valid scorer return type, but it carries only the number with no place to attach an explanation, so it cannot satisfy the requirement for a visible written rationale.
- C. A list return is meant for reporting several distinct named metrics from one scorer, not for pairing a single score with an explanation; it does not provide a structured rationale field either.
- D. A `str` can hold text, but jamming a score and reasoning into one unstructured sentence loses the distinct, queryable `value` and `rationale` fields that a `Feedback` object provides.
Sources2
4.When a scorer throws
If a scorer fails on one trace, MLflow records the error for that trace and carries on with the others. The recommended approach is to let exceptions propagate. MLflow catches the exception and creates a Feedback whose value is None and whose error holds the exception details, and you can open that row in the results to see them. When you want your own message, catch the exception and return a Feedback with error set. That can be the exception object itself, or an AssessmentError with a structured error_code.
if missing:
return Feedback(
error=AssessmentError(
error_code="MISSING_REQUIRED_FIELDS",
error_message=f"Missing required fields: {missing}",
),
)
return Feedback(
value=True,
rationale="Valid JSON with all required fields"
)
except json.JSONDecodeError as e:
return Feedback(error=e) # Can pass exception object directly to the error parameterCheckpoint 5 of 8· Check yourself
A custom scorer raises an uncaught KeyError on 3 of 200 traces. What does MLflow do?
Letting exceptions propagate is the recommended approach. MLflow turns each one into a Feedback with value None and the error attached.
“Let exceptions propagate (recommended) so that MLflow can capture error messages for you.”Source: docs.databricks.com
Sources3
5.@scorer vs the Scorer class, and running scorers in production
| Approach | Use when | Production monitoring |
|---|---|---|
| @scorer decorator | Most cases. Recommended starting point. | Supported (when defined and registered from a Databricks notebook) |
| Scorer class | You need stateful scorers, complex initialization, or Pydantic fields | Not supported |
The Scorer class is a Pydantic object. You must set its name field, add any extra fields you need, and put your logic in __call__. Keep state in instance attributes, not mutable class attributes: a class-level results = [] is shared by every instance.
# CORRECT: Use instance attributes
class GoodScorer(Scorer):
results: list[str] = None
name: str = "good_scorer"
def __init__(self):
self.results = [] # Per-instance state
def __call__(self, outputs, **kwargs):
self.results.append(outputs) # Safe
return Feedback(value=True)Production monitoring accepts built-in LLM judges and @scorer functions. It does not accept Scorer subclasses. If you need state in production, keep it inside the body of a decorated function. The function must also be defined and registered from a Databricks notebook, because the monitoring service serializes its code to run it remotely.
Scorers that call an external LLM endpoint can read credentials from Databricks secrets, both in development and in production. dbutils is not available in the scorer runtime by default, so import it inside the function.
Checkpoint 6 of 8· Fill the gap
Which module must the scorer import dbutils from to read secrets?
@scorer
def custom_llm_scorer(trace: Trace) -> Feedback:
# Explicitly import dbutils to access secrets
from ? import dbutils
# Retrieve your API key from Databricks secrets
api_key = dbutils.secrets.get(scope='my-scope', key='api-key')dbutils is not in the scorer runtime by default, so the documented pattern imports it from databricks.sdk.runtime inside the function.
Source: docs.databricks.comTo deploy a decorated scorer, register it and then start it with a sampling configuration. A sample_rate of 1 scores every trace.
# Register and start the scorer
custom_llm_scorer.register()
custom_llm_scorer.start(sampling_config = ScorerSamplingConfig(sample_rate=1))Checkpoint 7 of 8· Check yourself
A team built a stateful scorer as a Scorer subclass and now wants to run it on live traffic. What should they do?
Class-based scorers cannot be used in production monitoring. The documented route is a @scorer function, defined and registered from a Databricks notebook, that keeps its state internally.
“If you need stateful scorers in production, use the @scorer decorator and manage state inside the function body.”Source: docs.databricks.com
Checkpoint 8 of 8· Exam question
A team is writing a custom scorer that must flag agent runs where a particular tool span took longer than 2 seconds, using the full execution trace rather than just the final input and output text. Which parameter should the scorer function declare to access this span-level timing information?
Correct answer: A — Declare the `trace` parameter in the scorer function so the full set of spans, including tool-call timing, is available to inspect
- A. The `trace` parameter gives the scorer the complete MLflow trace, including every span, attribute, and timing detail, so inspecting individual tool-call durations requires declaring this parameter rather than relying on the final request or response text.
- B. `inputs` and `outputs` expose only the app's raw request and final response text; they do not carry span-level execution details like how long a specific tool call took, so they cannot support this check on their own.
- C. `expectations` holds ground truth such as an expected response or labeled facts supplied with the dataset; it is not a source of runtime span timing captured during execution.
- D. Access to the full trace, including span timing, comes from declaring the `trace` parameter in a function-based scorer; switching to the class-based approach does not add timing data that the decorator-based signature already exposes.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Any custom scorer that works in mlflow.genai.evaluate() can be registered for production monitoring.Why is that wrong?
Production monitoring supports built-in judges and @scorer functions registered from a Databricks notebook. Scorer subclasses are not supported.
Covered in @scorer vs the Scorer class, and running scorers in production
2.A scorer that depends on expectations works the same way in production monitoring as in offline evaluation.Why is that wrong?
Live traffic has no ground truth. Registered production scorers receive inputs and outputs only, so a reused scorer has to cope without expectations.
Covered in Expectations: available offline, missing in production
3.dbutils is available inside a scorer just as it is in a notebook cell, so secrets can be read directly.Why is that wrong?
The scorer runtime has no dbutils by default. You must import it from databricks.sdk.runtime inside the scorer function.
Covered in @scorer vs the Scorer class, and running scorers in production
Practise it for real
Develop a custom scorer against stored traces without rerunning your app on every change
1.Create an eval_dataset list whose rows each have an "inputs" dict, and a traced app function (decorated with @mlflow.trace)
Why: Evaluation needs input rows, and tracing records what the scorer will read
You should see: A list of rows plus a callable app
2.Run mlflow.genai.evaluate(data=eval_dataset, predict_fn=sample_app, scorers=[placeholder_metric]), where placeholder_metric is a @scorer that returns 1
Why: evaluate() needs at least one scorer, and this run exists only to produce traces
You should see: One trace in the experiment for each dataset row
3.Store the traces with generated_traces = mlflow.search_traces(run_id=eval_results.run_id)
Why: The stored traces become reusable evaluation input
You should see: A Pandas DataFrame of traces
4.Write a real @scorer, for example response_length on outputs, and run mlflow.genai.evaluate(data=generated_traces, scorers=[response_length]) without predict_fn
Why: Leaving out predict_fn scores the stored traces without calling the app
You should see: A response_length metric for each trace, with no new LLM calls from the app
Stuck? Get a nudge
If outputs is not a plain string for your app, declare trace instead and read the span you need.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/mlflow3/genai/eval-monitor/custom-scorer-dev-workflowOfficial docs
“Use this developer workflow to update your scorer without rerunning your entire app each time”
↩︎ Iterating on a scorer against stored traces“The mlflow.search_traces() function returns a Pandas DataFrame of traces.”
↩︎ Iterating on a scorer against stored traces“Since evaluate() requires at least one scorer, define a placeholder scorer for this initial trace generation”
↩︎ Prediction“This allows you to quickly iterate on your metric without having to re-run your app.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/mlflow3/genai/eval-monitor/code-based-scorer-examplesOfficial docs
“Access the full MLflow Trace object to use various details (spans, inputs, outputs, attributes, timing) for fine-grained metric calculation.”
↩︎ Two common patterns: reading spans and wrapping a judge“Use this to preprocess trace data for the judge or post-process its feedback.”
↩︎ Two common patterns: reading spans and wrapping a judge“design it to handle expectations gracefully”
↩︎ Expectations: available offline, missing in production“Production monitoring typically doesn't have expectations since you are evaluating live traffic without ground truth.”
↩︎ Exam trap 2 - 3.
“MLflow automatically captures the exception and creates a Feedback object with the following error details”
↩︎ When a scorer throws“AssessmentError: For structured error reporting with error codes.”
↩︎ When a scorer throws“be sure to use instance attributes, not mutable class attributes.”
↩︎ @scorer vs the Scorer class, and running scorers in production“This approach works for both development evaluation and production monitoring.”
↩︎ @scorer vs the Scorer class, and running scorers in production“By default, dbutils isn't available in the scorer runtime environment.”
↩︎ Exam trap 3“Registered scorers for production monitoring always parse the inputs and outputs parameters from the trace. expectations is not available.”
↩︎ Checkpoint“Let exceptions propagate (recommended) so that MLflow can capture error messages for you.”
↩︎ Checkpoint - 4.
“@scorer-decorated functions used in production monitoring must be defined and registered from a Databricks notebook.”
↩︎ @scorer vs the Scorer class, and running scorers in production“Class-based Scorer subclasses are not supported for production monitoring.”
↩︎ Exam trap 1“If you need stateful scorers in production, use the @scorer decorator and manage state inside the function body.”
↩︎ Checkpoint