What you will be able to do
- Explain why traces are the shared data layer behind evaluation, debugging and monitoring
- Read the mlflow.genai.evaluate() signature and say what each parameter is for
- Choose between direct evaluation and answer sheet evaluation for a given situation
- Retrieve evaluated traces and their judge feedback after an evaluation run
Key concept
Trace → scorer → Feedback — Every evaluation in MLflow starts from a trace, which is the recorded execution of your agent. A scorer reads that trace, judges one quality dimension, and attaches its verdict to the trace as Feedback. This works the same way whether evaluate() or the production monitoring service hands over the trace.
1.Traces are the evidence every score is based on
Agents run multi-step workflows. One request can pass through an LLM call, a retriever, a tool and a sub-agent before anything comes back. If you only see the final output, you can say that something went wrong, but not where. MLflow Tracing fills that gap. It records the inputs, outputs, latency, token usage and cost of every intermediate step, and arranges them as a tree of spans.
The documentation treats tracing as the base layer rather than a debugging extra. It calls tracing "the foundational data layer for agent quality" and says that debugging, evaluation and production monitoring all build on what traces capture. This is why the rest of this lesson keeps returning to traces. An evaluation run produces traces. Scorers read traces. Judge verdicts are stored on traces.
You can get traces with very little code. Autologging (for example mlflow.langchain.autolog() or mlflow.openai.autolog()) captures a trace with a span for each step. The docs' evaluation examples also decorate the app function with @mlflow.trace ("Your agent with MLflow tracing"), and that is the function later passed to evaluate() as predict_fn. To view a trace, open the experiment, click the Traces tab, and open a trace to see its span tree (agent → LLM call → tool call). In the trace panel you can also inspect each step's inputs, outputs, timing and errors.
Checkpoint 1 of 5· Check yourself
Which statement best describes the role of MLflow Tracing in agent quality work on Databricks?
Tracing captures every intermediate step, and the docs present it as the foundation that debugging, evaluation and production monitoring all rely on.
“Tracing records the inputs, outputs, latency, token usage, and cost of every intermediate step”Source: docs.databricks.com
It parses the trace to pull out the fields it needs. It runs its quality assessment on those fields. It returns the result as Feedback, which is attached to the trace.
2.The mlflow.genai.evaluate() harness
You could test an agent by hand: run it, read the output, and decide whether it was good. MLflow Evaluation replaces that with a structured run. You give it test data, it runs your agent, and it scores the results automatically, so you can compare versions and share results. The entry point is mlflow.genai.evaluate(), which runs the agent against an evaluation dataset with the scorers you choose and returns an EvaluationResult. The docs list typical uses: nightly or weekly checks against curated datasets, validating prompt or model changes, and checks before a release or PR.
def mlflow.genai.evaluate(
data: Union[pd.DataFrame, List[Dict], mlflow.genai.datasets.EvaluationDataset], # Test data.
scorers: list[mlflow.genai.scorers.Scorer], # Quality metrics, built-in or custom.
predict_fn: Optional[Callable[..., Any]] = None, # App wrapper. Used for direct evaluation only.
model_id: Optional[str] = None, # Optional version tracking.
) -> mlflow.models.evaluation.base.EvaluationResult:| Parameter | Required? | Role |
|---|---|---|
| data | Yes | Test data: a pandas DataFrame, a list of dicts, or an EvaluationDataset |
| scorers | Yes | The quality metrics to apply, built-in or custom |
| predict_fn | No (defaults to None) | Wraps your app; used only in direct evaluation |
| model_id | No | Optional version tracking |
What does predict_fn receive? In direct evaluation each row's inputs is a dictionary passed to your predict_fn, so the function is the entry point of your app, and the docs' example wraps it with @mlflow.trace so each call is captured as a trace. If the app is deployed as a Model Serving endpoint, you pass it wrapped in to_predict_fn instead. The model_id parameter is described in the docs only as optional version tracking.
Checkpoint 2 of 5· Check yourself
In the signature of mlflow.genai.evaluate(), when is the predict_fn parameter used?
predict_fn is optional and wraps your app for direct evaluation. Version tracking is the job of model_id, and answer sheet evaluation omits predict_fn.
“App wrapper. Used for direct evaluation only.”Source: docs.databricks.com
Two operational settings matter on real agents. MLflow evaluates rows in parallel on a background threadpool, and you set the number of workers with MLFLOW_GENAI_EVAL_MAX_WORKERS. If the model behind your agent is rate-limited, MLflow 3.11.1 and later can throttle predict_fn calls through MLFLOW_GENAI_EVAL_PREDICT_RATE_LIMIT.
data and scorers are required. predict_fn is optional and is used for direct evaluation only. model_id is optional version tracking.
Sources4
3.Direct evaluation vs. answer sheet evaluation
Direct evaluation is the recommended mode. You pass predict_fn, which is your app's entry point wrapped in a Python function. If the app is a Model Serving endpoint, you wrap it with to_predict_fn instead. MLflow then calls the app on each row's inputs, captures traces, and applies the scorers. Each row needs an inputs dict. An expectations dict holding ground truth is optional. The main reason to prefer this mode is that it produces the same kind of traces as production, so the scorers you write for offline evaluation can be reused unchanged for production monitoring.
# Evaluate your app
results = mlflow.genai.evaluate(
data=[
{"inputs": {"question": "What is MLflow?"}},
{"inputs": {"question": "How do I get started?"}}
],
predict_fn=my_chatbot_app,
scorers=[RelevanceToQuery(), Safety()]
)Answer sheet evaluation is for cases where you can't run the agent, or don't want to. Each row supplies either inputs plus pre-computed outputs, or an existing trace, with optional expectations in both cases. There is no predict_fn. When you give inputs and outputs, evaluate() builds traces from them first. A common pattern is to pull existing traces with mlflow.search_traces() and score them directly. One caution from the docs: if your answer-sheet traces look different from production traces, you may need to rewrite your scorers before you can use them for monitoring.
| Mode / input form | Required fields | Optional fields |
|---|---|---|
| Direct evaluation (with predict_fn) | inputs (dict passed to your predict_fn) | expectations |
| Answer sheet: inputs and outputs | inputs, outputs | expectations |
| Answer sheet: existing traces | trace (an mlflow.entities.Trace) | expectations |
import mlflow
# Retrieve traces from production
traces = mlflow.search_traces(
filter_string="trace.status = 'OK'",
)
# Evaluate problematic traces
evaluation = mlflow.genai.evaluate(
data=traces,
scorers=[Safety(), RelevanceToQuery()]
)Checkpoint 3 of 5· Check yourself
A team wants the scorers from its offline evaluation to run unchanged in production monitoring later. Which evaluation mode does Databricks recommend, and why?
Direct evaluation generates traces by calling the app, so they match production traces and the same scorers carry over. Both modes accept expectations and LLM judges.
“this mode enables you to reuse the scorers defined for offline evaluation in production monitoring since the resulting traces will be identical.”Source: docs.databricks.com
Sources4
4.Reading the results: traces annotated with feedback
An evaluation run does not produce a separate report. It produces traces with scorer feedback attached: one trace per dataset row, annotated by every judge. In the UI, open the experiment and click Evaluation runs. You get a table of traces with Pass/Fail assessments, and hovering over a label shows the judge's rationale. Click a request to open the full trace and step through each span's inputs and outputs. You can also add your own Feedback or Expectations there. In code, search for the run's traces. The result is a pandas DataFrame whose assessments column holds each judge's feedback.
Checkpoint 4 of 5· Fill the gap
Which attribute of the evaluation result identifies the run whose traces you want to retrieve?
eval_traces = mlflow.search_traces(run_id=eval_results. ? )Evaluation results are stored as traces linked to the run, so passing the result's run_id to search_traces returns that run's traces with their assessments.
Source: docs.databricks.comHow do you use spans to diagnose a failure? In the docs' RAG tutorial, the retrieval step is marked with span_type="RETRIEVER", which is what enables MLflow's retrieval-specific judges such as RetrievalGroundedness. The tutorial's improvement advice for RAG apps is to enhance retrieval "by examining retrieval spans" when relevant documents aren't being found, and to use prompt engineering to address failure patterns in the generation step. So a practical reading is: open the retriever span first and check whether it returned relevant context. If it did not, the problem is retrieval. If it did and the answer still failed, look at the LLM span and the prompt.
Sources5
5.Built-in judges and ground truth
The scorers you pass to evaluate() are often built-in LLM judges. These are predefined scorers that use Databricks-hosted LLMs to assess common quality dimensions such as relevance, safety, groundedness and correctness. Some judges are reference-free and need only the inputs and outputs. Others compare the response with ground truth, which you supply in each row's expectations field.
| Judge | Arguments | Requires ground truth |
|---|---|---|
| RelevanceToQuery | inputs, outputs | No |
| Safety | inputs, outputs | No |
| RetrievalGroundedness | inputs, outputs | No |
| Guidelines | inputs, outputs | No |
| Correctness | inputs, outputs, expectations | Yes |
| RetrievalSufficiency | inputs, outputs, expectations | Yes |
| ToolCallCorrectness | inputs, outputs, expectations | Yes |
Checkpoint 5 of 5· Exam question
An engineer has curated an evaluation dataset for a support FAQ agent, where each row includes an `expected_response` written by a subject-matter expert. They want to run `mlflow.genai.evaluate()` to score how factually accurate each generated answer is against that expected response. Which built-in judge should they register for this evaluation run?
Correct answer: A — Correctness, since it is a ground-truth-required judge that scores factual accuracy by comparing the generated response to the expected_response field
- A. This is correct because this judge is designed to take an expected_response field from the evaluation dataset and score factual accuracy of the generated answer against it, which is exactly the ground-truth comparison the engineer needs.
- B. This judge only checks topical relevance between the query and the response and does not consult expected_response, so it cannot measure factual accuracy against the SME-written answer.
- C. This judge screens for harmful or toxic content in the response and never compares against expected_response, so it does not measure factual accuracy at all.
- D. This judge evaluates adherence to custom style or format instructions supplied as natural-language rules, not against a ground-truth answer, so it would not score factual accuracy here.
Sources6
6.Code-based scorers with @scorer
When built-in and custom LLM judges don't fit, write a code-based scorer: a Python function you define for a custom heuristic, for mapping trace data to a built-in judge, or for using your own LLM. Most are defined with the @scorer decorator. The scorer can declare only the arguments it needs from inputs, outputs, expectations and trace. It can return a simple value such as a number, a boolean or a "yes"/"no" string, or a Feedback object with a rationale. For primitive returns, the function name becomes the metric name.
@scorer
def contains_citation(outputs: str) -> str:
# Return pass/fail string
return "yes" if "[source]" in outputs else "no"Sources7
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.mlflow.genai.evaluate() always needs a predict_fn, so you cannot score outputs you already have.Why is that wrong?
predict_fn is optional and used only for direct evaluation. In answer sheet mode you pass inputs with pre-computed outputs, or existing traces, and evaluate() runs only the scorers.
2.Tracing is a debugging tool, separate from evaluation and monitoring.Why is that wrong?
Traces are what evaluation and monitoring work on. Scorers read traces, and evaluation results are stored as traces with feedback.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Tracing is the foundational data layer for agent quality.”
↩︎ Traces are the evidence every score is based on“Open the trace to see the full span tree (agent → LLM call → tool call).”
↩︎ Traces are the evidence every score is based on“Tracing is the foundational data layer for agent quality.”
↩︎ Exam trap 2“Tracing records the inputs, outputs, latency, token usage, and cost of every intermediate step”
↩︎ Checkpoint - 2.
“Returns the quality assessment as Feedback to attach to the trace”
↩︎ Traces are the evidence every score is based on“A scorer receives a Trace from either evaluate() or the monitoring service.”
↩︎ Key concept - 3.
“the span tree, and each step's inputs, outputs, timing, and errors.”
↩︎ Traces are the evidence every score is based on - 4.
“The mlflow.genai.evaluate() function systematically tests GenAI agent quality by running it against test data (evaluation datasets) and applying scorers.”
↩︎ The mlflow.genai.evaluate() harness“Before a release or PR to prevent quality regressions”
↩︎ The mlflow.genai.evaluate() harness“inputs | dict[Any, Any] | Yes | Dictionary passed to your predict_fn”
↩︎ The mlflow.genai.evaluate() harness“Direct evaluation (recommended). MLflow calls your app directly to generate traces for evaluation”
↩︎ Direct evaluation vs. answer sheet evaluation“If you use an answer sheet with different traces than your production environment, you may need to re-write your scorer functions”
↩︎ Direct evaluation vs. answer sheet evaluation“If inputs and pre-computed outputs are provided, mlflow.genai.evaluate() constructs traces from the inputs and outputs.”
↩︎ Exam trap 1“App wrapper. Used for direct evaluation only.”
↩︎ Checkpoint“You provide the inputs and the output, and evaluate() runs scorers and logs an evaluation run.”
↩︎ Prediction“this mode enables you to reuse the scorers defined for offline evaluation in production monitoring since the resulting traces will be identical.”
↩︎ Checkpoint - 5.
“Running mlflow.genai.evaluate() creates traces with scorer feedback.”
↩︎ Reading the results: traces annotated with feedback“The column `assessments` includes each judge's feedback.”
↩︎ Reading the results: traces annotated with feedback“The retrieval component is marked with span_type="RETRIEVER" to enable MLflow's retrieval-specific LLM judges.”
↩︎ Reading the results: traces annotated with feedback“Enhance retrieval mechanisms if relevant documents aren't being found by examining retrieval spans”
↩︎ Reading the results: traces annotated with feedback - 6.
“Correctness | inputs, outputs, expectations | Yes | Is the response correct as compared to the provided ground truth?”
↩︎ Built-in judges and ground truth - 7.
“Most code-based scorers should be defined using the @scorer decorator.”
↩︎ Code-based scorers with @scorer