What you will be able to do
- Explain what an MLflow scorer receives, does and returns
- Choose between built-in judges, custom LLM judges, code-based scorers and third-party scorers
- Write a @scorer function that declares only the inputs it needs
- Predict which return type and metric name a scorer produces in the results
Key concept
Scorer — A scorer is an evaluation function. It takes an MLflow trace of one app interaction, pulls out the fields it needs, judges quality, and attaches the result to that trace as Feedback. A code-based custom scorer is a scorer whose logic you write yourself in Python.
1.What a scorer does with a trace
In the MLflow GenAI evaluation framework, scorers are the single interface for defining evaluation criteria for models, agents and applications. Every scorer follows the same three steps. It receives a Trace, from either mlflow.genai.evaluate() during development or the monitoring service in production. It parses the trace to extract the fields it needs. It runs its assessment and returns the result as Feedback attached to the trace.
The result does not have to be a number. It could be a pass/fail, true/false, numerical value, or a categorical value. Because scorers work on traces and not on a particular dataset format, you can use the same scorer to evaluate in development and to monitor in production. That keeps your evaluation consistent across the application lifecycle.
The trace holds every span: intermediate steps, tool calls and timing. So a scorer can measure things the final answer cannot show, such as latency or which tools the agent called.
Checkpoint 1 of 6· Check yourself
After a scorer parses a trace and runs its quality assessment, what does it do with the result?
The last of the scorer's three steps is returning the quality assessment as Feedback that is attached to the trace.
“Returns the quality assessment as Feedback to attach to the trace”Source: docs.databricks.com
Sources1
2.Where custom scorers sit among the scorer types
Databricks describes scorer types as a ladder. Each step adds customization and complexity. The recommended path is to start with built-in judges, move to custom LLM judges for domain-specific criteria, and write code-based scorers for deterministic business logic.
| Approach | Level of customization | Typical use |
|---|---|---|
| Built-in judges | Minimal (moderate for Guidelines judges) | Quick LLM evaluation with judges such as Correctness and RetrievalGroundedness |
| Custom judges | Full | Fully customized LLM judges that can return numerical scores, categories or boolean values |
| Code-based scorers | Full | Programmatic, deterministic checks: exact matching, format validation, performance metrics |
| Third-party scorers | Full | Specialized metrics from open-source evaluation frameworks |
Custom LLM judges let you write evaluation criteria in natural language. Code-based scorers are for cases where neither built-in nor custom LLM judges fit. The documentation names four reasons to write one:
- to define a custom heuristic or programmatic metric - to change how trace data is mapped to a built-in LLM judge - to use your own LLM in place of a Databricks-hosted judge - to get more flexibility and control than a custom LLM judge offers
Checkpoint 2 of 6· Check yourself
A team needs to verify that every agent response is valid JSON with three required fields, and the result must be deterministic. Which approach fits best?
Format validation with a deterministic result is exactly what code-based scorers are for. LLM judges add non-deterministic interpretation that this check doesn't need.
“Programmatic and deterministic scorers that evaluate things like exact matching, format validation, and performance metrics.”Source: docs.databricks.com
3.The @scorer signature and its four inputs
Most code-based scorers should be written as ordinary Python functions with the @scorer decorator. MLflow passes the complete trace, and it also pulls out the commonly needed parts and passes them as keyword-only named arguments:
from mlflow.genai.scorers import scorer
from typing import Optional, Any
from mlflow.entities import Feedback
@scorer
def my_custom_scorer(
*, # All arguments are keyword-only
inputs: Optional[dict[str, Any]], # App's raw input, a dictionary of input argument names and values
outputs: Optional[Any], # App's raw output
expectations: Optional[dict[str, Any]], # Ground truth, a dictionary of label names and values
trace: Optional[mlflow.entities.Trace] # Complete trace with all spans and metadata
) -> Union[int, float, bool, str, Feedback, List[Feedback]]:
# Your evaluation logic hereThe four inputs are:
- inputs: the request sent to the app
- outputs: the app's response
- expectations: ground truth or labels
- trace: the full trace, passed as an instantiated mlflow.entities.trace class, so you can analyse intermediate steps, latency and tool usage
When you run mlflow.genai.evaluate(), the first three can come from the data argument or be parsed from the trace. You only declare the arguments your logic uses. A word-count metric needs nothing but outputs.
Checkpoint 3 of 6· Check yourself
You write @scorer def response_length(outputs: str) -> int:. Which statement is correct?
MLflow passes each named argument only when the function declares it, so a scorer that uses only outputs is valid.
“All input arguments are optional, so declare only what your scorer needs”Source: docs.databricks.com
Sources3
4.Return types and how metrics get their names
For simple pass/fail or numeric results, return a primitive value. The MLflow UI shows each type differently:
@scorer
def response_length(outputs: str) -> int:
# Return a numeric metric
return len(outputs.split())
@scorer
def contains_citation(outputs: str) -> str:
# Return pass/fail string
return "yes" if "[source]" in outputs else "no"For a detailed assessment, return a Feedback object. It carries a value plus an optional rationale, source (an AssessmentSource such as HUMAN, CODE or LLM_JUDGE) and metadata. To score several aspects in one pass, return a list of Feedback objects. Each one should set name, and each name appears as a separate metric.
Checkpoint 4 of 6· Match them up
Match each scorer return type to how the MLflow UI shows it
Tap a term, then the definition that fits it.
Primitive values cover simple pass/fail or numeric results. Feedback adds a rationale, and a list of Feedback objects produces one metric per named entry.
“Return a Feedback object or list of Feedback objects for detailed assessments with scores, rationales, and metadata.”Source: docs.databricks.com
Names matter because they become metric names in evaluation results, monitoring results and dashboards. Use clear names such as safety_check or relevance_monitor. The rules are:
- A primitive value or an unnamed Feedback takes the function name with @scorer, or the Scorer.name field with the class.
- A named Feedback always uses its own name.
- All metrics must have distinct names, and that includes each Feedback in a returned list.
Checkpoint 5 of 6· Check yourself
A @scorer function named quality_gate returns Feedback(value=True) with no name. What is the metric called?
An unnamed Feedback from a decorated function takes the function name. A name is required only when you return a list, so the names stay distinct.
“For primitive return values or unnamed Feedbacks, the function name (for the @scorer decorator) or the Scorer.name field (for the Scorer class) is used.”Source: docs.databricks.com
Checkpoint 6 of 6· Exam question
An evaluation team needs a check that confirms each agent response contains a syntactically valid JSON object matching a required schema — a deterministic pass/fail check that does not need any LLM judgment. Which approach fits this requirement, and what should it return?
Correct answer: A — Define a function decorated with `@scorer` that reads `outputs` and returns a `bool` reporting whether the JSON matches the schema
- A. A deterministic structural check like schema validation is exactly what a code-based scorer is for: a function decorated with `@scorer` can parse `outputs` directly and return a `bool`, avoiding the cost and non-determinism of an LLM call for a check that has a clear right answer.
- B. Guidelines evaluates a response against a natural-language rule using an LLM judge, which introduces probabilistic judgment for something that has one objectively correct answer, so it is not the deterministic fit the team needs.
- C. Correctness requires ground truth in the form of an expected_response or expected_facts entry per case; it is not designed to validate JSON syntax or schema structure, and the team has not described a labeled reference dataset.
- D. Routing a purely structural, deterministic check through an LLM judge inside a custom scorer adds unnecessary cost and variability; a code-based `bool` check is more appropriate than a written critique for this pass/fail requirement.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The metric name of a @scorer is always the Python function name.Why is that wrong?
If the scorer returns Feedback objects with name set, those names win. The function name is used only for primitive values or unnamed Feedback.
Covered in Return types and how metrics get their names
2.Custom code-based scorers can only compute heuristics and cannot involve an LLM.Why is that wrong?
One documented use for a code-based scorer is calling your own LLM in place of a Databricks-hosted judge. Another is changing how trace data is mapped to a built-in judge.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“A scorer receives a Trace from either evaluate() or the monitoring service.”
↩︎ What a scorer does with a trace“This could be a pass/fail, true/false, numerical value, or a categorical value.”
↩︎ What a scorer does with a trace“create custom code-based scorers for deterministic business logic”
↩︎ Where custom scorers sit among the scorer types“Returns the quality assessment as Feedback to attach to the trace”
↩︎ Key concept“Using your own LLM (rather than a Databricks-hosted LLM judge) for evaluation.”
↩︎ Exam trap 2“Returns the quality assessment as Feedback to attach to the trace”
↩︎ Checkpoint“Programmatic and deterministic scorers that evaluate things like exact matching, format validation, and performance metrics.”
↩︎ Checkpoint - 2.
“Use them when built-in LLM judges and custom LLM judges don't fit your evaluation needs.”
↩︎ Where custom scorers sit among the scorer types - 3.
“Scorers receive the complete MLflow trace containing all spans, attributes, and outputs.”
↩︎ The @scorer signature and its four inputs“The trace is passed to the custom scorer as an instantiated mlflow.entities.trace class.”
↩︎ The @scorer signature and its four inputs“Each feedback should have the name field specified, and those names are displayed as separate metrics in the evaluation results.”
↩︎ Return types and how metrics get their names“For evaluation and monitoring, all metrics must have distinct names.”
↩︎ Return types and how metrics get their names“If the scorer returns one or more Feedback objects, then Feedback.name fields take precedence, if specified.”
↩︎ Exam trap 1“All input arguments are optional, so declare only what your scorer needs”
↩︎ Checkpoint“If the scorer returns one or more Feedback objects, then Feedback.name fields take precedence, if specified.”
↩︎ Prediction“Return a Feedback object or list of Feedback objects for detailed assessments with scores, rationales, and metadata.”
↩︎ Checkpoint“For primitive return values or unnamed Feedbacks, the function name (for the @scorer decorator) or the Scorer.name field (for the Scorer class) is used.”
↩︎ Checkpoint