What you will be able to do
- Explain why expert review of traces improves an agent, and pick which traces to send to experts
- Choose between the Review App Chat UI and labeling existing traces, and list what reviewers need for each
- Create a labeling session, add traces to it, and share it with domain experts
- Define labeling schemas and tell feedback schemas from expectation schemas, including the predefined schemas that built-in judges use
Key concept
Feedback vs. expectation — Subject-matter experts record two kinds of judgment on a trace. Feedback grades what the agent actually produced. An expectation records what the correct answer should have been, so it is ground truth that later evaluation can test against.
1.Why domain experts review traces
Automated judges and metrics only know what you tell them. A domain expert knows what a correct answer looks like for your business. Databricks puts it plainly: one of the most effective ways to improve your agent is to have domain experts review and label existing traces. The tool built for this is the MLflow Review App. It gives you a structured way to collect expert feedback on real interactions with your application.
The documentation lists three uses for the Review App. Experts help you understand what high-quality, correct responses look like for specific queries. They collect input to align LLM judges with your business requirements. And their labels help you create evaluation datasets from production traces. All three come back in the rest of this lesson.
Choosing traces is the first real decision. Focus on cases where a human adds something an automated judge cannot. That includes traces with ambiguous or borderline quality, edge cases not covered by automated judges, and examples where automated metrics disagree with expected quality. Add representative samples of different user interaction patterns so the labels aren't all edge cases. In the MLflow UI you can filter traces by status, tags, or time range. You can also select them through the SDK.
One note on direction: Databricks says that for new human-review workflows it recommends review queues (Beta), which route items to reviewers one at a time. The sources say little more about review queues, so this lesson covers labeling sessions and the Chat UI, which the docs describe in full.
Checkpoint 1 of 5· Check yourself
Which of these is NOT one of the documented uses of the Review App for labeling existing traces?
The documented uses are defining good responses, aligning judges, and building evaluation datasets. Nothing in the Review App retrains the model.
“Create evaluation datasets from production traces”Source: docs.databricks.com
2.Two ways to use the Review App: Chat UI or labeling existing traces
The Review App works in two ways, and they fit different situations.
The Chat UI lets experts talk to a live version of your agent and rate each answer as they go. Databricks calls it a way to vibe check your app. It is good for testing multi-turn conversations and for checking an update in a safe environment before production. You get it by calling get_review_app() and connecting your endpoint with review_app.add_agent(...). Then you share the URL with the experts. The Chat UI has fixed feedback questions, so you don't define a labeling schema. Every interaction and its feedback is automatically saved as a trace in MLflow.
Labeling existing traces is the more systematic way. You choose traces that are already logged, and experts answer questions that you define. It also covers deployments the Chat UI can't reach.
| Aspect | Chat UI | Label existing traces |
|---|---|---|
| What experts review | Their own queries sent to the live agent endpoint | Traces you already logged and selected |
| Deployment requirement | Agent deployed to a Model Serving endpoint (Databricks Apps not supported) | Works however the agent is deployed |
| Questions asked | Fixed feedback questions, no custom labeling schema | Labeling schemas you define |
| Reviewer permissions (beyond account provisioning and Workspace access) | CAN_QUERY on the model serving endpoint | Databricks SQL access entitlement and CAN_EDIT on the MLflow experiment |
| Minimum MLflow version | 3.1.0 | 3.14.0 |
Both modes need two things from every reviewer. They must be provisioned in your Databricks account, and they must have the Workspace access entitlement for the workspace that hosts the experiment. The docs warn that account provisioning alone is not enough. The other permissions depend on the mode, as the table shows.
Checkpoint 2 of 5· Check yourself
Your agent is deployed on Databricks Apps, and a compliance SME needs to review its answers. What should you set up?
The Chat UI only connects to Model Serving endpoints. Labeling existing traces works however the agent is deployed.
“Reviewing and labeling existing traces works regardless of how your agent is deployed.”Source: docs.databricks.com
3.Labeling sessions: packaging traces for expert review
To label existing traces, you organise them into a labeling session. A labeling session is a special type of MLflow run that holds the set of traces you want experts to review. Because it is an MLflow run, it shows up in the Evaluations tab of the MLflow UI, and you can query it with mlflow.search_runs().
When you create a session you set its name, the assigned users (the experts), an optional agent (this links the session to the Chat UI), the labeling schemas (the questions experts answer), and whether to allow multi-turn chat. Through the API, at least one schema is required.
import mlflow.genai.labeling as labeling
import mlflow.genai.label_schemas as schemas
# Create a simple labeling session with built-in schemas
session = labeling.create_labeling_session(
name="customer_service_review_jan_2024",
assigned_users=["alice@company.com", "bob@company.com"],
label_schemas=[schemas.EXPECTED_FACTS] # Required: at least one schema needed
)A new session is empty. After you create it, you must add traces for experts to review. In the UI, open the experiment's Traces view, tick the traces you want, and from the Actions drop-down menu select Add to labeling session. In code, use LabelingSession.add_traces. To bring in reviewers, open the session and click Share, then enter each reviewer's email address. Reviewers are notified and given access to the Review App. Their answers appear under Assessments when you open a request in the session.
One detail matters for automation. Session names might not be unique, so store and look up sessions by their MLflow run ID (session.mlflow_run_id), not by name.
Checkpoint 3 of 5· Put it in order
Put these steps for creating a labeling session in the MLflow UI in order
- 1.Click Create session
- 2.Enter a name, optionally select labeling schemas, and click Create Session
- 3.Click Labeling sessions in the sidebar
- 4.Click the name of your experiment to open it
- 5.In the left sidebar of the workspace, click Experiments
You go from the Experiments list into one experiment and then its Labeling sessions page. Only then does the Create Labeling Session dialog open, where you name the session and pick schemas.
“Enter a name for the session. You can also optionally specify an evaluation dataset or select labeling schemas.”Source: docs.databricks.com
Sources4
4.Labeling schemas: deciding what experts are asked
A labeling schema sets the questions an expert answers when labeling existing traces. It controls the question text, the input method (for example a drop-down or a text box), validation rules, and optional instructions. Each schema produces one assessment on the trace. Schemas are scoped to an experiment, so a name must be unique within that experiment. They are not used in the Chat UI.
Every schema has one of two types, matching the key concept of this lesson. A feedback schema collects subjective assessments such as ratings, preferences, or opinions. An expectation schema collects objective ground truth such as correct answers or expected behavior.
# Expectation schema for ground truth
facts_schema = schemas.create_label_schema(
name="required_facts",
type="expectation",
title="What facts must be included in a correct response?",
input=InputTextList(max_count=5, max_length_each=200),
instruction="List key facts that any correct response must contain."
)MLflow also ships predefined schema names that the built-in judges read. If you collect expert ground truth under one of these names, the matching built-in judge can use it without extra wiring. To change a schema later, call create_label_schema again with overwrite=True.
| Schema name | What experts provide | Used by built-in judges |
|---|---|---|
| GUIDELINES | Ideal instructions the agent should follow for a request | ExpectationGuidelines |
| EXPECTED_FACTS | Factual statements that must be included for correctness | Correctness, RetrievalSufficiency |
| EXPECTED_RESPONSE | The complete ground-truth answer | Correctness, RetrievalSufficiency |
Checkpoint 4 of 5· Match them up
Match each schema or schema type to what it collects
Tap a term, then the definition that fits it.
The three predefined names are all expectation schemas that built-in judges read. The feedback type is for subjective judgments about the output.
“Collects factual statements that must be included for correctness.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A GenAI engineering team wants to build a high-quality evaluation dataset from production traces by having subject matter experts (SMEs) supply correct reference answers that can later be used to test agent correctness. Which approach should they use?
Correct answer: A — Create a labeling session with `create_labeling_session()` and attach an Expectation-type schema (e.g. `EXPECTED_RESPONSE`) so SMEs record ground-truth answers per trace.
- A. Expectation-type schemas such as `EXPECTED_RESPONSE` are designed to capture objective ground-truth answers during a labeling session, which is exactly the material needed to build an evaluation dataset for testing correctness.
- B. A Feedback-type schema with a numeric scale captures a subjective quality rating, not a reference answer, so it does not give the team the ground truth needed to test correctness.
- C. The `GUIDELINES` schema captures general ideal instructions for the app rather than a per-trace ground-truth answer, and it bypasses the structured labeling session workflow entirely.
- D. Inference tables capture raw request and response traffic for monitoring, but they don't provide the Review App's schema-based workflow that SMEs use to record structured ground-truth labels.
Sources5
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The Review App Chat UI can connect to any deployed agent, including one running on Databricks Apps.Why is that wrong?
The Chat UI needs an agent deployed to a Model Serving endpoint. For Databricks Apps deployments, you label existing traces instead.
Covered in Two ways to use the Review App: Chat UI or labeling existing traces
2.You must define a custom labeling schema before experts can give feedback in the Chat UI.Why is that wrong?
Schemas apply only when labeling existing traces. The Chat UI uses fixed feedback questions.
Covered in Labeling schemas: deciding what experts are asked
3.A labeling session's name uniquely identifies it, so pipelines can look sessions up by name.Why is that wrong?
Session names can repeat. Use the MLflow run ID to store and reference a session.
Covered in Labeling sessions: packaging traces for expert review
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/mlflow3/genai/human-feedback/expert-feedback/label-existing-tracesOfficial docs
“One of the most effective ways to improve your agent is to have domain experts review and label existing traces.”
↩︎ Why domain experts review traces“Examples where automated metrics disagree with expected quality”
↩︎ Why domain experts review traces“Experiment access: CAN_EDIT permission on the MLflow experiment.”
↩︎ Two ways to use the Review App: Chat UI or labeling existing traces“Traces with ambiguous or borderline quality”
↩︎ Prediction“Create evaluation datasets from production traces”
↩︎ Checkpoint“because the Review App queries a SQL warehouse to display traces and record labels”
↩︎ Prediction - 2.
“For new human-review workflows, Databricks recommends review queues (Beta), which route traces and dataset records to reviewers one item at a time.”
↩︎ Why domain experts review traces“feedback evaluates the app's actual output (was the response good?), and an expectation defines the correct, ground-truth output.”
↩︎ Key concept - 3.https://docs.databricks.com/aws/en/mlflow3/genai/human-feedback/expert-feedback/live-app-testingOfficial docs
“Use the Chat UI as a way to vibe check your app.”
↩︎ Two ways to use the Review App: Chat UI or labeling existing traces“All interactions and feedback collected through the Chat UI are automatically captured as traces in MLflow.”
↩︎ Two ways to use the Review App: Chat UI or labeling existing traces“Endpoint access: CAN_QUERY permission on the model serving endpoint.”
↩︎ Two ways to use the Review App: Chat UI or labeling existing traces“The Review App Chat UI requires an agent deployed to a Model Serving endpoint. It does not currently support agents deployed on Databricks Apps.”
↩︎ Exam trap 1“You don't need to provide a custom labeling schema, as this approach uses fixed feedback questions.”
↩︎ Exam trap 2 - 4.https://docs.databricks.com/aws/en/mlflow3/genai/human-feedback/concepts/labeling-sessionsOfficial docs
“A labeling session is a special type of MLflow run that contains a specific set of traces that you want domain experts to review”
↩︎ Labeling sessions: packaging traces for expert review“After you create a session, you must add traces to it for expert review.”
↩︎ Labeling sessions: packaging traces for expert review“Enter an email address for each reviewer and click Save. Reviewers are notified and given access to the review app.”
↩︎ Labeling sessions: packaging traces for expert review“Session names might not be unique. Use the MLflow run ID (session.mlflow_run_id) to store and reference sessions.”
↩︎ Exam trap 3“Reviewing and labeling existing traces works regardless of how your agent is deployed.”
↩︎ Checkpoint“Enter a name for the session. You can also optionally specify an evaluation dataset or select labeling schemas.”
↩︎ Checkpoint - 5.https://docs.databricks.com/aws/en/mlflow3/genai/human-feedback/concepts/labeling-schemasOfficial docs
“They are not used for vibe checks in the Review App Chat UI.”
↩︎ Labeling schemas: deciding what experts are asked“expectation: Objective ground truth like correct answers or expected behavior.”
↩︎ Labeling schemas: deciding what experts are asked“MLflow provides predefined schema names for the built-in LLM Judges that use expectations.”
↩︎ Labeling schemas: deciding what experts are asked“Collects factual statements that must be included for correctness.”
↩︎ Checkpoint