What you will be able to do
- Turn an expert's ground-truth expectation on a trace into an evaluation dataset test case
- Describe the three-step judge alignment workflow and which judges it supports
- Log human feedback so alignment can use it, with the right name, volume, and balance
- Align a judge with align() and register the aligned version
1.From expert expectations to evaluation datasets
Collecting expert labels only pays off when they change how you evaluate the agent. Experts can record two kinds of judgment on a trace. Feedback grades the agent's actual output. An expectation records the correct, ground-truth output. Labeling sessions capture either kind, and the docs say this data can then be used to improve your agent through systematic evaluation.
Expectations do the most direct work. Databricks puts it this way: a trace that carries a ground-truth expectation becomes a ready-made test case for automated evaluation. You can promote it straight from the trace view, so the expert's correct answer becomes something every future version of the agent is tested against.
Checkpoint 1 of 6· Put it in order
Put these steps for turning a real interaction into an evaluation test case in order
- 1.Click Add to dataset to add the trace and its expectation to an evaluation dataset
- 2.Click Assess, then Add expectation to record the correct answer
- 3.In the Traces tab, click a trace to open it
You open the trace, record the ground-truth expectation on it, and then add the trace with its expectation to a dataset, where it becomes a test case.
“Click Add to dataset to add the trace — with its expectation — to an evaluation dataset.”Source: docs.databricks.com
The labels live on the trace, so the developer doesn't have to do this alone. Anyone who reviews traces can add an expectation and promote it without leaving the trace UI. Feedback from end users of the deployed app works alongside this. It surfaces problematic queries and preserves successful interactions, and those traces are good candidates for the next round of expert review.
2.Judge alignment: teaching LLM judges your experts' standards
The second use of expert feedback is less obvious. It fixes the judges that score your agent. A generic LLM judge may disagree with your experts about what counts as relevant or correct. Judge alignment teaches LLM judges to match human evaluation standards through systematic feedback. According to the docs, this improves agreement with human assessments by 30 to 50 percent compared to baseline judges.
Alignment works on built-in judges such as RelevanceToQuery, Safety, or Correctness, and on custom judges created with make_judge(). It needs MLflow 3.14.0 or above. There is one limit: alignment is not supported for session-level (multi-turn) judges such as ConversationCompleteness.
Checkpoint 2 of 6· Put it in order
Put the three steps of judge alignment in order
- 1.Collect human feedback: domain experts review and correct the judge's assessments
- 2.Generate initial assessments: run the judge on traces to establish a baseline
- 3.Align and deploy: call the judge's align() method to create a more aligned judge
The judge's baseline assessments come first, so experts have something to correct. Their corrections then become the training signal for align().
“Collect human feedback: Domain experts review and correct judge assessments.”Source: docs.databricks.com
Checkpoint 3 of 6· Exam question
A team's custom LLM judge for evaluating response correctness frequently disagrees with domain expert assessments on production traces. They want to systematically improve the judge's agreement with expert judgment without hand-tuning the judge's prompt. Which approach should they take?
Correct answer: A — Collect human feedback on traces where the feedback name matches the judge's `name`, then call the judge's `align()` method on those traces to produce an aligned judge.
- A. Calling `align()` on traces that carry human feedback matching the judge's name is the documented workflow for recalibrating a judge, and it can meaningfully raise agreement with expert assessments.
- B. Manually growing the few-shot examples in the prompt is ad hoc prompt engineering, not the systematic alignment workflow, and gives no guarantee of improved agreement with expert labels.
- C. Built-in reference-free judges like Safety check for a different quality dimension and are not automatically calibrated to a team's domain-specific correctness standards.
- D. Adjusting temperature changes the consistency of the judge's outputs across repeated runs, not the accuracy of its judgments relative to human labels.
Sources3
3.Logging expert feedback that alignment can use
Alignment connects the expert's label to the judge's assessment by name. The human feedback assessment name must exactly match the judge's name attribute. For a built-in judge, that is the default snake_case name, such as relevance_to_query for RelevanceToQuery, unless you passed name= when creating it. For a custom judge, it is the name you gave make_judge(), such as product_quality. If an expert labels under any other name, alignment cannot match that label to the judge's assessment.
Checkpoint 4 of 6· Check yourself
You want to align a default RelevanceToQuery judge. Under which name must your experts log their feedback?
Feedback must use exactly the judge's name attribute. For a default RelevanceToQuery instance, that is the snake_case relevance_to_query.
“For built-in judges, this is the default snake_case name (for example, relevance_to_query for RelevanceToQuery)”Source: docs.databricks.com
There are two ways to collect the feedback. In Databricks UI review, experts open the experiment's Traces tab, review each trace and the judge's assessment, and add feedback under the matching name. Use programmatic feedback when you already have ground-truth labels, a large set of examples, or need collection to be reproducible.
| Approach | Use it when |
|---|---|
| Databricks UI review | You need domain experts to review outputs, want to iteratively refine feedback criteria, or have a smaller dataset (< 100 examples) |
| Programmatic feedback | You have pre-existing ground truth labels, large datasets (100+ examples), or need reproducible feedback collection |
# Log human feedback for each trace
for item in ground_truth_data:
mlflow.log_feedback(
trace_id=item["trace_id"],
name=initial_judge.name, # Must match judge name (built-in or custom)
value=item["label"],
rationale=item.get("rationale", ""),
source=AssessmentSource(
source_type=AssessmentSourceType.HUMAN,
source_id="ground_truth_dataset"
),
)How much data, and of what kind? You can get reasonable alignment from at least 10 traces, but 50–100 traces give better results. The quality guidance is specific. Use diverse reviewers, meaning several domain experts. Keep the examples balanced, with at least 30% negative examples (poor or fair ratings). Ask for clear rationales that explain each rating. And choose representative samples that cover edge cases as well as common scenarios. A set made only of good answers teaches the judge almost nothing about what failure looks like.
It falls far short of the guidance to include at least 30% negative examples (poor or fair ratings). With almost no failures to learn from, the aligned judge can't learn where your experts draw the line.
Sources3
4.Running align() and registering the aligned judge
Once enough traces carry both the judge's assessment and an expert's assessment, retrieve them and call align(). Built-in and custom judges use the same method. If you call align() without naming an optimizer, MLflow uses the MemAlign optimizer by default. Other optimizers are available in the package mlflow.genai.judges.optimizers.
Checkpoint 5 of 6· Fill the gap
Which method completes this alignment sample?
if len(traces_for_alignment) >= 10:
# Align the judge based on human feedback using the default optimizer
aligned_judge = initial_judge. ? (traces_for_alignment)align() takes traces that carry both judge and human assessments and returns a new aligned judge. register() comes afterwards, to save that judge for production use.
Source: docs.databricks.comThe aligned judge is a new object. Register it for production use under a new name so you can tell it apart from the original, and tag it with details such as how many traces it was aligned on. From then on, your evaluations and monitoring score the agent with a judge that reflects your experts' standards, not generic criteria.
aligned_judge.register(
experiment_id=experiment_id,
name=f"{initial_judge.name}_aligned",
tags={"alignment_date": "2025-10-23", "num_traces": str(len(traces_for_alignment))}
)Checkpoint 6 of 6· Exam question
An organization contracts an external subject matter expert to review and label agent traces in the Review App as part of a labeling session, but the expert should not receive broader access to the Databricks workspace. What is the minimum access the expert needs?
Correct answer: A — CAN_EDIT permission on the MLflow experiment that contains the traces, plus account-level access provisioned for that expert.
- A. Reviewing traces through the Review App requires only account-level access plus CAN_EDIT on the specific MLflow experiment; no separate workspace access is required for the SME.
- B. Granting manage rights on the whole workspace is far broader than what a Review App reviewer needs and exposes notebooks and compute unrelated to labeling.
- C. A metastore-scoped token lets someone query underlying tables directly, but it does not grant access to the Review App's labeling interface.
- D. Admin rights on the tracking server let a user create and delete experiments, which is far more privilege than a reviewer needs just to submit labels.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Any human feedback on a trace is used by align(), whatever name it was logged under.Why is that wrong?
Alignment only pairs feedback whose assessment name exactly matches the judge's name attribute.
2.Every judge can be aligned, including multi-turn judges such as ConversationCompleteness.Why is that wrong?
Alignment does not support session-level judges. It applies to built-in and custom judges that score individual traces.
Covered in Judge alignment: teaching LLM judges your experts' standards
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“a trace that carries a ground-truth expectation becomes a ready-made test case for automated evaluation.”
↩︎ From expert expectations to evaluation datasets“Capture feedback from users of your deployed app to surface problematic queries and preserve successful interactions.”
↩︎ From expert expectations to evaluation datasets“Click Add to dataset to add the trace — with its expectation — to an evaluation dataset.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/mlflow3/genai/human-feedback/concepts/labeling-sessionsOfficial docs
“You can capture either Feedback or Expectation data, which can then be used to improve your agent through systematic evaluation.”
↩︎ From expert expectations to evaluation datasets - 3.
“Judge alignment teaches LLM judges to match human evaluation standards through systematic feedback.”
↩︎ Judge alignment: teaching LLM judges your experts' standards“improving agreement with human assessments by 30 to 50 percent compared to baseline judges.”
↩︎ Judge alignment: teaching LLM judges your experts' standards“The same alignment workflow applies to both built-in judges (such as RelevanceToQuery, Safety, or Correctness) and custom judges created with make_judge().”
↩︎ Judge alignment: teaching LLM judges your experts' standards“You can achieve reasonable alignment with at least 10 traces, but 50-100 traces yield better results.”
↩︎ Logging expert feedback that alignment can use“Balanced examples: Include at least 30% negative examples (poor/fair ratings)”
↩︎ Logging expert feedback that alignment can use“When you call align() without specifying an optimizer, the MemAlign optimizer is used automatically”
↩︎ Running align() and registering the aligned judge“The system supports the optimizers that are available in the package mlflow.genai.judges.optimizers.”
↩︎ Running align() and registering the aligned judge“The human feedback assessment name must exactly match the judge's name attribute.”
↩︎ Exam trap 1“Alignment is not supported for session-level (multi-turn) judges such as ConversationCompleteness.”
↩︎ Exam trap 2“Align and deploy: Invoke the judge's align() method to create a new judge that is more aligned with human feedback.”
↩︎ Prediction“Collect human feedback: Domain experts review and correct judge assessments.”
↩︎ Checkpoint“For built-in judges, this is the default snake_case name (for example, relevance_to_query for RelevanceToQuery)”
↩︎ Checkpoint