CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 25/56

    Tracing and Evaluating Agents with MLflow 3

    Utilize MLflow and Agent Framework for developing agentic systems

    12 min read
    1.79% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what an MLflow trace records and why debugging, evaluation and monitoring all depend on it
    • Turn on automatic tracing for an agent framework and know when autologging must be called explicitly
    • Run mlflow.genai.evaluate() in direct mode or answer-sheet mode and pick the right one
    • Describe how scorers and LLM judges produce Feedback, and how the same scorers carry over to production monitoring

    1.Why tracing comes first

    An agent's run can include LLM calls, retrievers, tools and sub-agents. MLflow Tracing records each of those intermediate steps as part of the trace, with its inputs, outputs, latency, token usage and cost. Debugging, evaluation and production monitoring all work from what the traces capture. On Databricks you can store traces as Unity Catalog Delta tables, which gives you schema-level and table-level access control and lets you query them directly with SQL.

    Checkpoint 1 of 7· Check yourself

    Which statement about MLflow traces on Databricks is correct?

    Sources12

    2.Turning on tracing with autolog

    MLflow supports automatic tracing for more than 30 frameworks and LLM providers. Calling the framework's autolog function patches the library, so each LLM invocation, tool call and agent step is captured as a span without any other instrumentation. LangGraph has no autolog of its own and uses LangChain's. On serverless compute autologging is not on by default, so you must call mlflow.<library>.autolog() explicitly for each integration you want traced. Point the tracking URI at Databricks, set an experiment, and run the agent. The trace then appears in that experiment. When autolog doesn't cover part of your code, manual tracing lets you add custom spans; the evaluation example later in this lesson wraps its app function with the @mlflow.trace decorator.

    A traced LangGraph agent: one autolog call plus an experiment, and every step becomes a spanpython
    import mlflow
    from langchain_core.tools import tool
    from langchain_openai import ChatOpenAI
    from langgraph.prebuilt import create_react_agent
    
    mlflow.langchain.autolog()  # LangGraph uses LangChain's autolog
    mlflow.set_tracking_uri("databricks")
    mlflow.set_experiment("/Shared/my-first-trace")
    
    @tool
    def get_weather(city: str):
        """Get weather for a city."""
        return f"It might be cloudy in {city}"
    
    agent = create_react_agent(ChatOpenAI(model="gpt-4o-mini"), [get_weather])
    agent.invoke({"messages": [("user", "What is the weather in SF?")]})
    # Trace appears in the MLflow UI automatically

    Checkpoint 2 of 7· Fill the gap

    This code calls Databricks Foundation Model APIs through the OpenAI client. Which autolog integration does it enable?

    # Databricks Foundation Model APIs use the OpenAI client
    mlflow. ? .autolog()
    mlflow.set_tracking_uri("databricks")
    mlflow.set_experiment("/Shared/databricks-fmapi-tracing")

    Checkpoint 3 of 7· Exam question

    An engineering team is iterating on a multi-step LangChain agent that calls several tools before producing a final answer. They want every LLM call and tool invocation automatically captured as MLflow trace spans during local development, without adding manual `@mlflow.trace` decorators to each function. What should they do?

    Sources31

    3.Evaluating with mlflow.genai.evaluate()

    You don't need to run the agent and check its outputs one at a time. mlflow.genai.evaluate() takes test data and scorers, optionally a predict function or model ID, and returns an EvaluationResult. The documentation suggests running it nightly or weekly against curated datasets, when you validate prompt or model changes, and before a release or PR to catch quality regressions. It uses a background threadpool, and you set the number of workers with MLFLOW_GENAI_EVAL_MAX_WORKERS.

    The evaluate() signature: predict_fn is used only in direct evaluationpython
    def mlflow.genai.evaluate(
        data: Union[pd.DataFrame, List[Dict], mlflow.genai.datasets.EvaluationDataset],  # Test data.
        scorers: list[mlflow.genai.scorers.Scorer],  # Quality metrics, built-in or custom.
        predict_fn: Optional[Callable[..., Any]] = None,  # App wrapper. Used for direct evaluation only.
        model_id: Optional[str] = None,  # Optional version tracking.
    ) -> mlflow.models.evaluation.base.EvaluationResult:
    The two evaluation modes compared
    AspectDirect evaluation (recommended)Answer sheet evaluation
    Who runs the agentMLflow calls your predict_fn, or a serving endpoint wrapped in to_predict_fnNobody; you supply results already produced
    Required data fieldsinputs (expectations optional)inputs and outputs, or trace (expectations optional)
    Reuse scorers in production monitoringYes, because the resulting traces are identicalScorers may need rewriting if the traces differ from production

    Answer-sheet mode fits when you already have outputs, for example from external systems, historical traces or batch jobs. One common pattern is to pull production traces and score them.

    Answer-sheet evaluation over traces pulled from productionpython
    import mlflow
    
    # Retrieve traces from production
    traces = mlflow.search_traces(
        filter_string="trace.status = 'OK'",
    )
    
    # Evaluate problematic traces
    evaluation = mlflow.genai.evaluate(
        data=traces,
        scorers=[Safety(), RelevanceToQuery()]
    )

    Checkpoint 4 of 7· Check yourself

    A nightly batch job has already produced responses for 5,000 questions, and you only want to score them. What do you pass to evaluate()?

    Checkpoint 5 of 7· Exam question

    A developer is prototyping a retrieval tool locally in a notebook before deploying an agent that queries a Databricks-managed vector search index built with Databricks-managed embeddings. They want a LangChain-compatible tool object they can pass directly to their agent with minimal setup. Which option fits this stage of development?

    Sources4

    4.Scorers and LLM judges

    A scorer receives a Trace, either from evaluate() or from the monitoring service. It pulls the fields it needs out of the trace, runs its assessment, and returns the result as Feedback attached to that trace. LLM judges are scorers that use an LLM to do the assessment. MLflow includes built-in judges for relevance, safety, groundedness and correctness, plus multi-turn judges that assess whole conversations. You can also write custom judges, for example when you need graded scores rather than pass/fail, or need to check that the agent made the right decisions. A scorer does not have to use an LLM: a code-based check, such as the custom exact_match scorer shown alongside the built-in Safety judge in the documentation, is also a scorer. Judges can be used directly with evaluate() or wrapped in custom scorers for more advanced scoring logic. By default a judge uses a Databricks-hosted LLM, and you can swap it with the model argument.

    Swapping the judge model with the <provider>:/<model-name> formatpython
    from mlflow.genai.scorers import Correctness
    
    Correctness(model="databricks:/databricks-gpt-5-mini")

    Checkpoint 6 of 7· Check yourself

    What does a scorer produce once it has assessed a trace?

    Sources52

    5.The improvement loop, from trace to production

    Databricks describes agent improvement as a loop you keep running, where each pass fixes one issue. Monitoring uses the same LLM judges and custom metrics as offline evaluation, so a scorer written to catch a bug in development keeps checking live traffic for the same bug. That is the main reason direct evaluation is recommended: its traces match production traces.

    Checkpoint 7 of 7· Put it in order

    Put the stages of the agent quality loop in order

    1. 1.Trace your agent
    2. 2.Collect feedback and curate an evaluation dataset
    3. 3.Write a scorer that catches the issue
    4. 4.Re-run evaluation to confirm the fix
    5. 5.Find an issue in development traces or production
    6. 6.Fix the agent

    Sources62

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Tracing is automatic on all Databricks compute, so you never need to call autolog.Why is that wrong?

      On serverless compute you have to call autolog explicitly for each integration you want traced.

      Covered in Turning on tracing with autolog

    2. 2.Scorers built for answer-sheet evaluation always work unchanged in production monitoring.Why is that wrong?

      Only direct evaluation guarantees traces identical to production. If the answer sheet's traces differ, you may have to rewrite the scorers.

      Covered in Evaluating with mlflow.genai.evaluate()

    Practise it for real

    Trace a small LangGraph agent and score a traced app with built-in judges

    1. 1.In a Databricks notebook, run %pip install --upgrade "mlflow[databricks]>=3.1.0" "langgraph" "langchain-openai" and then dbutils.library.restartPython(). Then either set OPENAI_API_KEY, or point the agent at Databricks Foundation Model APIs to avoid an external key

      Why: MLflow 3.1+ provides the GenAI tracing and evaluation APIs, and the agent needs a model to call

      You should see: The libraries install, Python restarts, and the agent has a model it can call

    2. 2.Run the traced LangGraph agent from this lesson, which calls mlflow.langchain.autolog() and mlflow.set_experiment("/Shared/my-first-trace") before agent.invoke

      Why: autolog captures the agent, LLM and tool steps as spans, and autologging is not on by default on serverless

      You should see: The invocation returns a message about the weather in SF

    3. 3.Go to Experiments, open /Shared/my-first-trace, click the Traces tab and open the trace

      Why: Every later step depends on these traces

      You should see: A span tree of agent, then LLM call, then tool call

    4. 4.Call mlflow.set_experiment() to choose an experiment for the evaluation run. Then wrap a function with @mlflow.trace and pass it as predict_fn to mlflow.genai.evaluate() with scorers=[RelevanceToQuery(), Safety()] and two {"inputs": {"question": ...}} records. Afterwards open that experiment's Traces tab

      Why: Direct evaluation stores results as traces with scorer feedback in the active experiment, and its traces match production, so these scorers can later be reused for monitoring

      You should see: An evaluation run whose traces, visible in the experiment, carry Feedback from each scorer

    Stuck? Get a nudge

    If no trace appears, check that autolog ran before the agent was created and that the experiment path matches.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Debugging, evaluation, and production monitoring all build on the evidence your traces capture.”
      ↩︎ Why tracing comes first
      “Add custom spans when autolog doesn't cover your code.”
      ↩︎ Turning on tracing with autolog
      “Tracing records the inputs, outputs, latency, token usage, and cost of every intermediate step”
      ↩︎ Checkpoint
    2. 2.
      “With only an input and final response, you cannot tell which intermediate decision caused a failure.”
      ↩︎ Why tracing comes first
      “Define a scorer (an LLM judge or code-based check) that catches the issue automatically.”
      ↩︎ Scorers and LLM judges
      “Each pass fixes one issue and surfaces the next.”
      ↩︎ The improvement loop, from trace to production
      “Trace your agent. Instrument your agent so every execution is captured as a trace”
      ↩︎ Checkpoint
    3. 3.
      “It patches the library at import time so every LLM invocation, tool call, and agent step is captured as a span”
      ↩︎ Turning on tracing with autolog
      “On serverless compute clusters, autologging is not enabled by default.”
      ↩︎ Exam trap 1
      “On serverless compute clusters, autologging is not enabled by default.”
      ↩︎ Prediction
    4. 4.
      “this mode enables you to reuse the scorers defined for offline evaluation in production monitoring”
      ↩︎ Evaluating with mlflow.genai.evaluate()
      “Before a release or PR to prevent quality regressions”
      ↩︎ Evaluating with mlflow.genai.evaluate()
      “you may need to re-write your scorer functions to use them for production monitoring.”
      ↩︎ Exam trap 2
      “you already have outputs (for example, from external systems, historical traces, or batch jobs) and you just want to score them.”
      ↩︎ Checkpoint
    5. 5.
      “LLM judges are a type of MLflow Scorer that uses Large Language Models for quality assessment.”
      ↩︎ Scorers and LLM judges
      “Returns the quality assessment as Feedback to attach to the trace”
      ↩︎ Checkpoint
    6. 6.
      “Use the same evaluation configuration (LLM judges and custom metrics) in offline evaluation and online monitoring.”
      ↩︎ The improvement loop, from trace to production

    Spotted a mistake, or was something unclear? Tell us.