What you will be able to do
- Choose between built-in judges, custom judges, code-based scorers and third-party scorers
- Identify which built-in judges need ground truth (expectations) and which do not
- Write Guidelines judges and change the model a judge uses
- Use evaluation results to improve an agent and check that nothing regressed
1.Four kinds of scorer, from least to most control
A scorer returns a judgement on some aspect of quality. The judgement can be pass/fail, true/false, a number or a category. Databricks presents the scorer options as a ladder. Start with built-in judges, add custom LLM judges for criteria specific to your domain, and write code-based scorers for deterministic business logic. LLM judges are scorers that use an LLM to make the assessment. They can look at inputs, outputs and the whole execution trace, and they recognise when two differently worded queries mean the same thing. The same scorer object works in development evaluation and in production monitoring, which keeps quality measured the same way across the lifecycle.
| Approach | Customization | Typical use |
|---|---|---|
| Built-in judges | Minimal (moderate for Guidelines judges) | Fast LLM evaluation, e.g. Correctness, RetrievalGroundedness |
| Custom judges | Full | Domain-specific LLM criteria returning numbers, categories or booleans |
| Code-based scorers | Full | Deterministic checks such as exact matching, format validation, performance metrics |
| Third-party scorers | Full | Specialized metrics from open-source evaluation frameworks |
Checkpoint 1 of 3· Check yourself
Every agent response must be valid JSON with a fixed set of keys. Which scorer approach fits best?
Format validation is deterministic logic, which is exactly what code-based scorers are for. An LLM judge would add cost and variability without any benefit.
“Programmatic and deterministic scorers that evaluate things like exact matching, format validation, and performance metrics.”Source: docs.databricks.com
Sources1
2.Built-in judges and the ground-truth question
| Judge | Requires ground truth | What it evaluates |
|---|---|---|
| RelevanceToQuery | No | Response is relevant to the user's request |
| RetrievalRelevance | No | Retrieved context is relevant to the request |
| Safety | No | Content is free of harmful or toxic material |
| RetrievalGroundedness | No | Response is grounded in retrieved context (hallucination) |
| Guidelines | No | Response meets specified natural-language criteria |
| ToolCallEfficiency | No | Tool calls are free of redundancy |
| ExpectationsGuidelines | No (but needs guidelines in expectations) | Per-example natural-language criteria |
| Correctness | Yes | Response is correct compared with ground truth |
| RetrievalSufficiency | Yes | Context contains everything needed for the ground-truth facts |
| ToolCallCorrectness | Yes | Tool calls and arguments are correct for the query |
The pattern is that a judge needs ground truth when it compares against a known right answer: Correctness, RetrievalSufficiency and ToolCallCorrectness. Reference-free judges can therefore score unlabeled traffic, while ground-truth judges need an evaluation dataset that has expectations filled in. This matters for production monitoring: registered monitoring scorers always parse inputs and outputs from the trace, and expectations is not available to them. So the ground-truth judges cannot get their labels in monitoring, and reference-free judges are the ones that fit live traffic. For conversational agents there are also multi-turn judges that take a whole session, such as ConversationCompleteness, UserFrustration and KnowledgeRetention. None of these needs ground truth.
The retrieval judges also depend on how you instrument the app. They need to find the retriever step in the trace, so the Databricks tutorial marks its retrieval function with span_type="RETRIEVER".
Checkpoint 2 of 3· Match them up
Match each situation to the built-in judge that fits it
Tap a term, then the definition that fits it.
Safety and ToolCallEfficiency work on inputs and outputs alone, so they suit unlabeled traffic. Correctness and RetrievalSufficiency both require ground truth in expectations.
“Is the response correct as compared to the provided ground truth?”Source: docs.databricks.com
3.Guidelines, judge models and code-based scorers
Guidelines judges are the quickest way to encode your own rules. Each one is a named, natural-language rule that the judge marks pass or fail. In the Databricks email-agent tutorial, RetrievalGroundedness, RelevanceToQuery and Safety run alongside five Guidelines judges, for instruction following, conciseness, naming the contact, professional tone and concrete next steps. All of them are passed together as scorers to mlflow.genai.evaluate().
Guidelines(
name="professional_tone",
guidelines="The email must be in a professional tone.",
),By default, each judge runs on a Databricks-hosted LLM designed for quality assessment. To use a different model, pass the model argument in the form <provider>:/<model-name>.
from mlflow.genai.scorers import Correctness
Correctness(model="databricks:/databricks-gpt-5-mini")When neither of those is enough, there are two further options. Custom LLM judges give you control over grades and scores beyond pass/fail. They can also check whether the agent made the right decisions, and judge alignment can tune them to match human evaluation standards. Custom code-based scorers are defined with the @scorer decorator or the Scorer class. Use them for custom heuristics, for changing how trace data is mapped to a built-in judge, or for running evaluation on your own LLM instead of a Databricks-hosted judge.
Most code-based scorers use the @scorer decorator. A scorer receives the complete trace, and MLflow also passes commonly needed data as named, keyword-only arguments: inputs (the request sent to the app), outputs (the app's response), expectations (ground truth or labels) and trace (all spans, for analysing intermediate steps, latency or tool use). All of them are optional, so declare only what your scorer needs. In mlflow.genai.evaluate(), inputs, outputs and expectations can come from the data argument or be parsed from the trace.
@scorer
def contains_citation(outputs: str) -> str:
# Return pass/fail string
return "yes" if "[source]" in outputs else "no"| Return Type | MLflow UI Display | Use Case |
|---|---|---|
| "yes"/"no" | Pass/Fail | Binary evaluation |
| True/False | True/False | Boolean checks |
| int/float | Numeric value | Scores, counts |
| Feedback | Value + rationale | Detailed assessment |
| List[Feedback] | Multiple metrics | Multi-aspect evaluation |
To run a scorer on live traffic, register it with the experiment and then start it with a sampling configuration. This .register() then .start() pattern works for built-in judges, custom judges, code-based scorers and multi-turn judges. sample_rate sets the fraction of traces the scorer evaluates. Databricks recommends sample_rate=1.0 for critical scorers such as safety and security checks, and lower rates (0.05-0.2) for expensive scorers such as complex LLM judges. Two limits apply. Registered monitoring scorers parse inputs and outputs from the trace and have no expectations. Only @scorer-decorated functions, defined and registered from a Databricks notebook, are supported: class-based Scorer subclasses are not.
from mlflow.genai.scorers import Safety, ScorerSamplingConfig
# Register and start a built-in judge
safety_judge = Safety().register(name="safety")
safety_judge = safety_judge.start(sampling_config=ScorerSamplingConfig(sample_rate=0.7))4.From scores to a better agent
Scores only matter if they lead to changes. In the tutorial's first run, the feedback shows three failure patterns: weak instruction following, emails that are too long, and missing concrete next steps. The fixes target those patterns directly. Options include prompt engineering, guardrails that check outputs before users see them, retrieval improvements found by examining retrieval spans, and splitting reasoning into more spans. The judge list is saved in a variable, so the improved version is evaluated with exactly the same scorers. That makes the before-and-after comparison fair.
Checkpoint 3 of 3· Put it in order
Put the tutorial's evaluate-and-improve workflow in order
- 1.Interpret results to identify quality issues
- 2.Evaluate quality with LLM judges using the evaluation harness
- 3.Improve the app based on evaluation results
- 4.Compare versions to verify improvements and catch regressions
- 5.Create evaluation datasets from real usage data
The tutorial goes from building the dataset, to the first evaluation, to interpreting results, to a targeted fix, and finally to comparing versions with the same judges.
“Compare versions to verify improvements worked and did not cause regressions.”Source: docs.databricks.com
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Hallucination checks such as RetrievalGroundedness need a labeled expected answer.Why is that wrong?
RetrievalGroundedness compares the response with the retrieved context and needs only inputs and outputs. It is Correctness that compares against ground truth.
2.Any built-in judge can run on unlabeled data.Why is that wrong?
Correctness, RetrievalSufficiency and ToolCallCorrectness all require ground truth in expectations, so the dataset must provide it.
3.A scorer registered for production monitoring can read expectations, so Correctness works on live traffic.Why is that wrong?
Registered monitoring scorers parse inputs and outputs from the trace and have no expectations, so ground-truth judges cannot get their labels there.
Practise it for real
Score pre-computed agent answers with built-in judges, then inspect the feedback on the resulting traces
1.Install mlflow[databricks]>=3.1.0 and set up an MLflow experiment
Why: Evaluation results are stored as traces in the active experiment
You should see: An experiment you can open under Experiments
2.Build a list of dicts, each with an inputs question and a pre-computed outputs response
Why: This is answer sheet evaluation, so no predict_fn is needed
You should see: Two or more rows with inputs and outputs keys
3.Call mlflow.genai.evaluate(data=..., scorers=[Safety(), RelevanceToQuery()])
Why: Both judges are reference-free, so the data needs no expectations
You should see: An EvaluationResult with a run_id
4.Call mlflow.search_traces(run_id=<result>.run_id) and print the assessments column
Why: Judge feedback is attached to each trace
You should see: One trace per row, each with Safety and RelevanceToQuery feedback
Stuck? Get a nudge
Try adding Correctness() without expectations, and notice that it needs ground truth that your rows do not contain.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“You can use the same scorer for evaluation in development and monitoring in production to keep evaluation consistent throughout the application lifecycle.”
↩︎ Four kinds of scorer, from least to most control“Specify the model in the format <provider>:/<model-name>.”
↩︎ Guidelines, judge models and code-based scorers“Customizing how the data from your app's trace is mapped to built-in LLM judges.”
↩︎ Guidelines, judge models and code-based scorers“Programmatic and deterministic scorers that evaluate things like exact matching, format validation, and performance metrics.”
↩︎ Checkpoint - 2.
“Is the response correct as compared to the provided ground truth?”
↩︎ Built-in judges and the ground-truth question“Correctness | inputs, outputs, expectations | Yes”
↩︎ Exam trap 1“Does the context provide all necessary information to generate a response that includes the ground truth facts?”
↩︎ Exam trap 2 - 3.
“Registered scorers for production monitoring always parse the inputs and outputs parameters from the trace. expectations is not available.”
↩︎ Built-in judges and the ground-truth question“Define custom code-based scorers in MLflow using the @scorer decorator or the Scorer class.”
↩︎ Guidelines, judge models and code-based scorers“All input arguments are optional, so declare only what your scorer needs:”
↩︎ Guidelines, judge models and code-based scorers“Registered scorers for production monitoring always parse the inputs and outputs parameters from the trace. expectations is not available.”
↩︎ Exam trap 3 - 4.
“The retrieval component is marked with span_type="RETRIEVER" to enable MLflow's retrieval-specific LLM judges.”
↩︎ Built-in judges and the ground-truth question“Retrieval improvements (for RAG apps): Enhance retrieval mechanisms if relevant documents aren't being found by examining retrieval spans”
↩︎ From scores to a better agent“Compare versions to verify improvements worked and did not cause regressions.”
↩︎ Checkpoint - 5.
“For critical scorers such as safety and security checks, use sample_rate=1.0.”
↩︎ Guidelines, judge models and code-based scorers