What you will be able to do
- Explain what an MLflow trace records and why debugging, evaluation and monitoring all depend on it
- Turn on automatic tracing for an agent framework and know when autologging must be called explicitly
- Run mlflow.genai.evaluate() in direct mode or answer-sheet mode and pick the right one
- Describe how scorers and LLM judges produce Feedback, and how the same scorers carry over to production monitoring
1.Why tracing comes first
An agent's run can include LLM calls, retrievers, tools and sub-agents. MLflow Tracing records each of those intermediate steps as part of the trace, with its inputs, outputs, latency, token usage and cost. Debugging, evaluation and production monitoring all work from what the traces capture. On Databricks you can store traces as Unity Catalog Delta tables, which gives you schema-level and table-level access control and lets you query them directly with SQL.
You can't tell which intermediate step caused the failure: the retrieval, a tool call or the model's reasoning. A trace records every step, so you can find the root cause, collect feedback on that run, and add the failure to an evaluation dataset.
Checkpoint 1 of 7· Check yourself
Which statement about MLflow traces on Databricks is correct?
Tracing records every step. The same traces power debugging, evaluation and monitoring on fully managed MLflow.
“Tracing records the inputs, outputs, latency, token usage, and cost of every intermediate step”Source: docs.databricks.com
2.Turning on tracing with autolog
MLflow supports automatic tracing for more than 30 frameworks and LLM providers. Calling the framework's autolog function patches the library, so each LLM invocation, tool call and agent step is captured as a span without any other instrumentation. LangGraph has no autolog of its own and uses LangChain's. On serverless compute autologging is not on by default, so you must call mlflow.<library>.autolog() explicitly for each integration you want traced. Point the tracking URI at Databricks, set an experiment, and run the agent. The trace then appears in that experiment. When autolog doesn't cover part of your code, manual tracing lets you add custom spans; the evaluation example later in this lesson wraps its app function with the @mlflow.trace decorator.
import mlflow
from langchain_core.tools import tool
from langchain_openai import ChatOpenAI
from langgraph.prebuilt import create_react_agent
mlflow.langchain.autolog() # LangGraph uses LangChain's autolog
mlflow.set_tracking_uri("databricks")
mlflow.set_experiment("/Shared/my-first-trace")
@tool
def get_weather(city: str):
"""Get weather for a city."""
return f"It might be cloudy in {city}"
agent = create_react_agent(ChatOpenAI(model="gpt-4o-mini"), [get_weather])
agent.invoke({"messages": [("user", "What is the weather in SF?")]})
# Trace appears in the MLflow UI automaticallyCheckpoint 2 of 7· Fill the gap
This code calls Databricks Foundation Model APIs through the OpenAI client. Which autolog integration does it enable?
# Databricks Foundation Model APIs use the OpenAI client
mlflow. ? .autolog()
mlflow.set_tracking_uri("databricks")
mlflow.set_experiment("/Shared/databricks-fmapi-tracing")The Foundation Model APIs are called here with the OpenAI client, so the OpenAI autolog integration is what captures the trace.
Source: docs.databricks.comCheckpoint 3 of 7· Exam question
An engineering team is iterating on a multi-step LangChain agent that calls several tools before producing a final answer. They want every LLM call and tool invocation automatically captured as MLflow trace spans during local development, without adding manual `@mlflow.trace` decorators to each function. What should they do?
Correct answer: A — Call `mlflow.langchain.autolog()` at the start of the notebook so agent executions are automatically instrumented and logged as traces.
- A. This is correct because calling this autolog function once enables automatic instrumentation of LangChain agent execution, capturing LLM calls and tool invocations as trace spans without any per-function decorators.
- B. This is incorrect because wrapping each call in a run and logging metrics manually is exactly the manual instrumentation the team wants to avoid, and metrics do not produce the structured trace spans that tracing provides.
- C. This is incorrect because the autologging call is meant to be invoked once at the top of the notebook or script, not embedded inside a specific class definition, and doing so would not reliably instrument all agent calls.
- D. This is incorrect because exporting endpoint access logs captures raw request/response traffic, not the structured, hierarchical trace spans that MLflow tracing produces for agent tool calls and reasoning steps.
3.Evaluating with mlflow.genai.evaluate()
You don't need to run the agent and check its outputs one at a time. mlflow.genai.evaluate() takes test data and scorers, optionally a predict function or model ID, and returns an EvaluationResult. The documentation suggests running it nightly or weekly against curated datasets, when you validate prompt or model changes, and before a release or PR to catch quality regressions. It uses a background threadpool, and you set the number of workers with MLFLOW_GENAI_EVAL_MAX_WORKERS.
def mlflow.genai.evaluate(
data: Union[pd.DataFrame, List[Dict], mlflow.genai.datasets.EvaluationDataset], # Test data.
scorers: list[mlflow.genai.scorers.Scorer], # Quality metrics, built-in or custom.
predict_fn: Optional[Callable[..., Any]] = None, # App wrapper. Used for direct evaluation only.
model_id: Optional[str] = None, # Optional version tracking.
) -> mlflow.models.evaluation.base.EvaluationResult:| Aspect | Direct evaluation (recommended) | Answer sheet evaluation |
|---|---|---|
| Who runs the agent | MLflow calls your predict_fn, or a serving endpoint wrapped in to_predict_fn | Nobody; you supply results already produced |
| Required data fields | inputs (expectations optional) | inputs and outputs, or trace (expectations optional) |
| Reuse scorers in production monitoring | Yes, because the resulting traces are identical | Scorers may need rewriting if the traces differ from production |
Answer-sheet mode fits when you already have outputs, for example from external systems, historical traces or batch jobs. One common pattern is to pull production traces and score them.
import mlflow
# Retrieve traces from production
traces = mlflow.search_traces(
filter_string="trace.status = 'OK'",
)
# Evaluate problematic traces
evaluation = mlflow.genai.evaluate(
data=traces,
scorers=[Safety(), RelevanceToQuery()]
)Checkpoint 4 of 7· Check yourself
A nightly batch job has already produced responses for 5,000 questions, and you only want to score them. What do you pass to evaluate()?
This is answer-sheet evaluation. Both inputs and outputs are required, and evaluate() builds traces from them before it runs the scorers.
“you already have outputs (for example, from external systems, historical traces, or batch jobs) and you just want to score them.”Source: docs.databricks.com
Checkpoint 5 of 7· Exam question
A developer is prototyping a retrieval tool locally in a notebook before deploying an agent that queries a Databricks-managed vector search index built with Databricks-managed embeddings. They want a LangChain-compatible tool object they can pass directly to their agent with minimal setup. Which option fits this stage of development?
Correct answer: A — Install `databricks-langchain` and instantiate `VectorSearchRetrieverTool` with the `index_name` and `tool_description` parameters.
- A. This is correct because this class from the Databricks AI Bridge package is purpose-built to wrap a Databricks-managed vector index as a ready-to-use LangChain tool with minimal setup for local prototyping.
- B. This is incorrect because an embeddings class produces vector representations of text; it is not itself a retriever tool and cannot be passed to an agent as a callable tool.
- C. This is incorrect because standing up a managed MCP server is the recommended path for production-grade, governed access, but it adds infrastructure overhead unnecessary for quick local prototyping against an already Databricks-managed index.
- D. This is incorrect because the Unity Catalog connection pattern is intended for indexes or vector stores hosted outside Databricks, not for an index that already uses Databricks-managed embeddings and hosting.
Sources4
4.Scorers and LLM judges
A scorer receives a Trace, either from evaluate() or from the monitoring service. It pulls the fields it needs out of the trace, runs its assessment, and returns the result as Feedback attached to that trace. LLM judges are scorers that use an LLM to do the assessment. MLflow includes built-in judges for relevance, safety, groundedness and correctness, plus multi-turn judges that assess whole conversations. You can also write custom judges, for example when you need graded scores rather than pass/fail, or need to check that the agent made the right decisions. A scorer does not have to use an LLM: a code-based check, such as the custom exact_match scorer shown alongside the built-in Safety judge in the documentation, is also a scorer. Judges can be used directly with evaluate() or wrapped in custom scorers for more advanced scoring logic. By default a judge uses a Databricks-hosted LLM, and you can swap it with the model argument.
from mlflow.genai.scorers import Correctness
Correctness(model="databricks:/databricks-gpt-5-mini")Checkpoint 6 of 7· Check yourself
What does a scorer produce once it has assessed a trace?
A scorer's result is Feedback attached to the trace it evaluated, so the assessment stays linked to that run.
“Returns the quality assessment as Feedback to attach to the trace”Source: docs.databricks.com
5.The improvement loop, from trace to production
Databricks describes agent improvement as a loop you keep running, where each pass fixes one issue. Monitoring uses the same LLM judges and custom metrics as offline evaluation, so a scorer written to catch a bug in development keeps checking live traffic for the same bug. That is the main reason direct evaluation is recommended: its traces match production traces.
Checkpoint 7 of 7· Put it in order
Put the stages of the agent quality loop in order
- 1.Trace your agent
- 2.Collect feedback and curate an evaluation dataset
- 3.Write a scorer that catches the issue
- 4.Re-run evaluation to confirm the fix
- 5.Find an issue in development traces or production
- 6.Fix the agent
Traces come first because every later step uses them. After the fix is evaluated, monitoring runs the scorers on live traffic.
“Trace your agent. Instrument your agent so every execution is captured as a trace”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Tracing is automatic on all Databricks compute, so you never need to call autolog.Why is that wrong?
On serverless compute you have to call autolog explicitly for each integration you want traced.
Covered in Turning on tracing with autolog
2.Scorers built for answer-sheet evaluation always work unchanged in production monitoring.Why is that wrong?
Only direct evaluation guarantees traces identical to production. If the answer sheet's traces differ, you may have to rewrite the scorers.
Covered in Evaluating with mlflow.genai.evaluate()
Practise it for real
Trace a small LangGraph agent and score a traced app with built-in judges
1.In a Databricks notebook, run %pip install --upgrade "mlflow[databricks]>=3.1.0" "langgraph" "langchain-openai" and then dbutils.library.restartPython(). Then either set OPENAI_API_KEY, or point the agent at Databricks Foundation Model APIs to avoid an external key
Why: MLflow 3.1+ provides the GenAI tracing and evaluation APIs, and the agent needs a model to call
You should see: The libraries install, Python restarts, and the agent has a model it can call
2.Run the traced LangGraph agent from this lesson, which calls mlflow.langchain.autolog() and mlflow.set_experiment("/Shared/my-first-trace") before agent.invoke
Why: autolog captures the agent, LLM and tool steps as spans, and autologging is not on by default on serverless
You should see: The invocation returns a message about the weather in SF
3.Go to Experiments, open /Shared/my-first-trace, click the Traces tab and open the trace
Why: Every later step depends on these traces
You should see: A span tree of agent, then LLM call, then tool call
4.Call mlflow.set_experiment() to choose an experiment for the evaluation run. Then wrap a function with @mlflow.trace and pass it as predict_fn to mlflow.genai.evaluate() with scorers=[RelevanceToQuery(), Safety()] and two {"inputs": {"question": ...}} records. Afterwards open that experiment's Traces tab
Why: Direct evaluation stores results as traces with scorer feedback in the active experiment, and its traces match production, so these scorers can later be reused for monitoring
You should see: An evaluation run whose traces, visible in the experiment, carry Feedback from each scorer
Stuck? Get a nudge
If no trace appears, check that autolog ran before the agent was created and that the experiment path matches.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Debugging, evaluation, and production monitoring all build on the evidence your traces capture.”
↩︎ Why tracing comes first“Add custom spans when autolog doesn't cover your code.”
↩︎ Turning on tracing with autolog“Tracing records the inputs, outputs, latency, token usage, and cost of every intermediate step”
↩︎ Checkpoint - 2.
“With only an input and final response, you cannot tell which intermediate decision caused a failure.”
↩︎ Why tracing comes first“Define a scorer (an LLM judge or code-based check) that catches the issue automatically.”
↩︎ Scorers and LLM judges“Each pass fixes one issue and surfaces the next.”
↩︎ The improvement loop, from trace to production“Trace your agent. Instrument your agent so every execution is captured as a trace”
↩︎ Checkpoint - 3.
“It patches the library at import time so every LLM invocation, tool call, and agent step is captured as a span”
↩︎ Turning on tracing with autolog“On serverless compute clusters, autologging is not enabled by default.”
↩︎ Exam trap 1“On serverless compute clusters, autologging is not enabled by default.”
↩︎ Prediction - 4.
“this mode enables you to reuse the scorers defined for offline evaluation in production monitoring”
↩︎ Evaluating with mlflow.genai.evaluate()“Before a release or PR to prevent quality regressions”
↩︎ Evaluating with mlflow.genai.evaluate()“you may need to re-write your scorer functions to use them for production monitoring.”
↩︎ Exam trap 2“you already have outputs (for example, from external systems, historical traces, or batch jobs) and you just want to score them.”
↩︎ Checkpoint - 5.
“LLM judges are a type of MLflow Scorer that uses Large Language Models for quality assessment.”
↩︎ Scorers and LLM judges“Returns the quality assessment as Feedback to attach to the trace”
↩︎ Checkpoint - 6.
“Use the same evaluation configuration (LLM judges and custom metrics) in offline evaluation and online monitoring.”
↩︎ The improvement loop, from trace to production