CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 49/56

    MLflow Tracing and the mlflow.genai.evaluate() Harness

    Evaluate agent performance using MLflow scoring and tracing

    13 min read
    1.79% of exam
    7 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why traces are the shared data layer behind evaluation, debugging and monitoring
    • Read the mlflow.genai.evaluate() signature and say what each parameter is for
    • Choose between direct evaluation and answer sheet evaluation for a given situation
    • Retrieve evaluated traces and their judge feedback after an evaluation run

    Key concept

    Trace → scorer → Feedback — Every evaluation in MLflow starts from a trace, which is the recorded execution of your agent. A scorer reads that trace, judges one quality dimension, and attaches its verdict to the trace as Feedback. This works the same way whether evaluate() or the production monitoring service hands over the trace.

    1.Traces are the evidence every score is based on

    Agents run multi-step workflows. One request can pass through an LLM call, a retriever, a tool and a sub-agent before anything comes back. If you only see the final output, you can say that something went wrong, but not where. MLflow Tracing fills that gap. It records the inputs, outputs, latency, token usage and cost of every intermediate step, and arranges them as a tree of spans.

    The documentation treats tracing as the base layer rather than a debugging extra. It calls tracing "the foundational data layer for agent quality" and says that debugging, evaluation and production monitoring all build on what traces capture. This is why the rest of this lesson keeps returning to traces. An evaluation run produces traces. Scorers read traces. Judge verdicts are stored on traces.

    You can get traces with very little code. Autologging (for example mlflow.langchain.autolog() or mlflow.openai.autolog()) captures a trace with a span for each step. The docs' evaluation examples also decorate the app function with @mlflow.trace ("Your agent with MLflow tracing"), and that is the function later passed to evaluate() as predict_fn. To view a trace, open the experiment, click the Traces tab, and open a trace to see its span tree (agent → LLM call → tool call). In the trace panel you can also inspect each step's inputs, outputs, timing and errors.

    Checkpoint 1 of 5· Check yourself

    Which statement best describes the role of MLflow Tracing in agent quality work on Databricks?

    Sources123

    2.The mlflow.genai.evaluate() harness

    You could test an agent by hand: run it, read the output, and decide whether it was good. MLflow Evaluation replaces that with a structured run. You give it test data, it runs your agent, and it scores the results automatically, so you can compare versions and share results. The entry point is mlflow.genai.evaluate(), which runs the agent against an evaluation dataset with the scorers you choose and returns an EvaluationResult. The docs list typical uses: nightly or weekly checks against curated datasets, validating prompt or model changes, and checks before a release or PR.

    The signature of mlflow.genai.evaluate()python
    def mlflow.genai.evaluate(
        data: Union[pd.DataFrame, List[Dict], mlflow.genai.datasets.EvaluationDataset],  # Test data.
        scorers: list[mlflow.genai.scorers.Scorer],  # Quality metrics, built-in or custom.
        predict_fn: Optional[Callable[..., Any]] = None,  # App wrapper. Used for direct evaluation only.
        model_id: Optional[str] = None,  # Optional version tracking.
    ) -> mlflow.models.evaluation.base.EvaluationResult:
    What each evaluate() parameter controls
    ParameterRequired?Role
    dataYesTest data: a pandas DataFrame, a list of dicts, or an EvaluationDataset
    scorersYesThe quality metrics to apply, built-in or custom
    predict_fnNo (defaults to None)Wraps your app; used only in direct evaluation
    model_idNoOptional version tracking

    What does predict_fn receive? In direct evaluation each row's inputs is a dictionary passed to your predict_fn, so the function is the entry point of your app, and the docs' example wraps it with @mlflow.trace so each call is captured as a trace. If the app is deployed as a Model Serving endpoint, you pass it wrapped in to_predict_fn instead. The model_id parameter is described in the docs only as optional version tracking.

    Checkpoint 2 of 5· Check yourself

    In the signature of mlflow.genai.evaluate(), when is the predict_fn parameter used?

    Two operational settings matter on real agents. MLflow evaluates rows in parallel on a background threadpool, and you set the number of workers with MLFLOW_GENAI_EVAL_MAX_WORKERS. If the model behind your agent is rate-limited, MLflow 3.11.1 and later can throttle predict_fn calls through MLFLOW_GENAI_EVAL_PREDICT_RATE_LIMIT.

    Sources4

    3.Direct evaluation vs. answer sheet evaluation

    Direct evaluation is the recommended mode. You pass predict_fn, which is your app's entry point wrapped in a Python function. If the app is a Model Serving endpoint, you wrap it with to_predict_fn instead. MLflow then calls the app on each row's inputs, captures traces, and applies the scorers. Each row needs an inputs dict. An expectations dict holding ground truth is optional. The main reason to prefer this mode is that it produces the same kind of traces as production, so the scorers you write for offline evaluation can be reused unchanged for production monitoring.

    Direct evaluation: MLflow calls my_chatbot_app (a @mlflow.trace-decorated function) on each inputpython
    # Evaluate your app
    results = mlflow.genai.evaluate(
        data=[
            {"inputs": {"question": "What is MLflow?"}},
            {"inputs": {"question": "How do I get started?"}}
        ],
        predict_fn=my_chatbot_app,
        scorers=[RelevanceToQuery(), Safety()]
    )

    Answer sheet evaluation is for cases where you can't run the agent, or don't want to. Each row supplies either inputs plus pre-computed outputs, or an existing trace, with optional expectations in both cases. There is no predict_fn. When you give inputs and outputs, evaluate() builds traces from them first. A common pattern is to pull existing traces with mlflow.search_traces() and score them directly. One caution from the docs: if your answer-sheet traces look different from production traces, you may need to rewrite your scorers before you can use them for monitoring.

    Row schemas the docs list for each mode (inputs+outputs and trace are separate schemas in answer sheet evaluation)
    Mode / input formRequired fieldsOptional fields
    Direct evaluation (with predict_fn)inputs (dict passed to your predict_fn)expectations
    Answer sheet: inputs and outputsinputs, outputsexpectations
    Answer sheet: existing tracestrace (an mlflow.entities.Trace)expectations
    Answer sheet evaluation over traces that already existpython
    import mlflow
    
    # Retrieve traces from production
    traces = mlflow.search_traces(
        filter_string="trace.status = 'OK'",
    )
    
    # Evaluate problematic traces
    evaluation = mlflow.genai.evaluate(
        data=traces,
        scorers=[Safety(), RelevanceToQuery()]
    )

    Checkpoint 3 of 5· Check yourself

    A team wants the scorers from its offline evaluation to run unchanged in production monitoring later. Which evaluation mode does Databricks recommend, and why?

    Sources4

    4.Reading the results: traces annotated with feedback

    An evaluation run does not produce a separate report. It produces traces with scorer feedback attached: one trace per dataset row, annotated by every judge. In the UI, open the experiment and click Evaluation runs. You get a table of traces with Pass/Fail assessments, and hovering over a label shows the judge's rationale. Click a request to open the full trace and step through each span's inputs and outputs. You can also add your own Feedback or Expectations there. In code, search for the run's traces. The result is a pandas DataFrame whose assessments column holds each judge's feedback.

    Checkpoint 4 of 5· Fill the gap

    Which attribute of the evaluation result identifies the run whose traces you want to retrieve?

    eval_traces = mlflow.search_traces(run_id=eval_results. ? )

    How do you use spans to diagnose a failure? In the docs' RAG tutorial, the retrieval step is marked with span_type="RETRIEVER", which is what enables MLflow's retrieval-specific judges such as RetrievalGroundedness. The tutorial's improvement advice for RAG apps is to enhance retrieval "by examining retrieval spans" when relevant documents aren't being found, and to use prompt engineering to address failure patterns in the generation step. So a practical reading is: open the retriever span first and check whether it returned relevant context. If it did not, the problem is retrieval. If it did and the answer still failed, look at the LLM span and the prompt.

    Sources5

    5.Built-in judges and ground truth

    The scorers you pass to evaluate() are often built-in LLM judges. These are predefined scorers that use Databricks-hosted LLMs to assess common quality dimensions such as relevance, safety, groundedness and correctness. Some judges are reference-free and need only the inputs and outputs. Others compare the response with ground truth, which you supply in each row's expectations field.

    Selected built-in judges and whether they need ground truth
    JudgeArgumentsRequires ground truth
    RelevanceToQueryinputs, outputsNo
    Safetyinputs, outputsNo
    RetrievalGroundednessinputs, outputsNo
    Guidelinesinputs, outputsNo
    Correctnessinputs, outputs, expectationsYes
    RetrievalSufficiencyinputs, outputs, expectationsYes
    ToolCallCorrectnessinputs, outputs, expectationsYes

    Checkpoint 5 of 5· Exam question

    An engineer has curated an evaluation dataset for a support FAQ agent, where each row includes an `expected_response` written by a subject-matter expert. They want to run `mlflow.genai.evaluate()` to score how factually accurate each generated answer is against that expected response. Which built-in judge should they register for this evaluation run?

    Sources6

    6.Code-based scorers with @scorer

    When built-in and custom LLM judges don't fit, write a code-based scorer: a Python function you define for a custom heuristic, for mapping trace data to a built-in judge, or for using your own LLM. Most are defined with the @scorer decorator. The scorer can declare only the arguments it needs from inputs, outputs, expectations and trace. It can return a simple value such as a number, a boolean or a "yes"/"no" string, or a Feedback object with a rationale. For primitive returns, the function name becomes the metric name.

    A simple code-based scorer returning a pass/fail stringpython
    @scorer
    def contains_citation(outputs: str) -> str:
        # Return pass/fail string
        return "yes" if "[source]" in outputs else "no"

    Sources7

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.mlflow.genai.evaluate() always needs a predict_fn, so you cannot score outputs you already have.Why is that wrong?

      predict_fn is optional and used only for direct evaluation. In answer sheet mode you pass inputs with pre-computed outputs, or existing traces, and evaluate() runs only the scorers.

      Covered in Direct evaluation vs. answer sheet evaluation

    2. 2.Tracing is a debugging tool, separate from evaluation and monitoring.Why is that wrong?

      Traces are what evaluation and monitoring work on. Scorers read traces, and evaluation results are stored as traces with feedback.

      Covered in Traces are the evidence every score is based on

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Tracing is the foundational data layer for agent quality.”
      ↩︎ Traces are the evidence every score is based on
      “Open the trace to see the full span tree (agent → LLM call → tool call).”
      ↩︎ Traces are the evidence every score is based on
      “Tracing is the foundational data layer for agent quality.”
      ↩︎ Exam trap 2
      “Tracing records the inputs, outputs, latency, token usage, and cost of every intermediate step”
      ↩︎ Checkpoint
    2. 2.
      “Returns the quality assessment as Feedback to attach to the trace”
      ↩︎ Traces are the evidence every score is based on
      “A scorer receives a Trace from either evaluate() or the monitoring service.”
      ↩︎ Key concept
    3. 3.
    4. 4.
      “The mlflow.genai.evaluate() function systematically tests GenAI agent quality by running it against test data (evaluation datasets) and applying scorers.”
      ↩︎ The mlflow.genai.evaluate() harness
      “Before a release or PR to prevent quality regressions”
      ↩︎ The mlflow.genai.evaluate() harness
      “inputs | dict[Any, Any] | Yes | Dictionary passed to your predict_fn”
      ↩︎ The mlflow.genai.evaluate() harness
      “Direct evaluation (recommended). MLflow calls your app directly to generate traces for evaluation”
      ↩︎ Direct evaluation vs. answer sheet evaluation
      “If you use an answer sheet with different traces than your production environment, you may need to re-write your scorer functions”
      ↩︎ Direct evaluation vs. answer sheet evaluation
      “If inputs and pre-computed outputs are provided, mlflow.genai.evaluate() constructs traces from the inputs and outputs.”
      ↩︎ Exam trap 1
      “App wrapper. Used for direct evaluation only.”
      ↩︎ Checkpoint
      “You provide the inputs and the output, and evaluate() runs scorers and logs an evaluation run.”
      ↩︎ Prediction
      “this mode enables you to reuse the scorers defined for offline evaluation in production monitoring since the resulting traces will be identical.”
      ↩︎ Checkpoint
    5. 5.
      “Running mlflow.genai.evaluate() creates traces with scorer feedback.”
      ↩︎ Reading the results: traces annotated with feedback
      “The column `assessments` includes each judge's feedback.”
      ↩︎ Reading the results: traces annotated with feedback
      “The retrieval component is marked with span_type="RETRIEVER" to enable MLflow's retrieval-specific LLM judges.”
      ↩︎ Reading the results: traces annotated with feedback
      “Enhance retrieval mechanisms if relevant documents aren't being found by examining retrieval spans”
      ↩︎ Reading the results: traces annotated with feedback
    6. 6.
      “Correctness | inputs, outputs, expectations | Yes | Is the response correct as compared to the provided ground truth?”
      ↩︎ Built-in judges and ground truth
    7. 7.
      “Most code-based scorers should be defined using the @scorer decorator.”
      ↩︎ Code-based scorers with @scorer

    Continue to page 2 of 2

    MLflow Scorers and LLM Judges for Agent Evaluation

    Spotted a mistake, or was something unclear? Tell us.