CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 56/56

    Use SME Labels to Build Evaluation Datasets and Align LLM Judges

    Incorporate SME feedback to improve agent performance

    9 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Turn an expert's ground-truth expectation on a trace into an evaluation dataset test case
    • Describe the three-step judge alignment workflow and which judges it supports
    • Log human feedback so alignment can use it, with the right name, volume, and balance
    • Align a judge with align() and register the aligned version

    1.From expert expectations to evaluation datasets

    Collecting expert labels only pays off when they change how you evaluate the agent. Experts can record two kinds of judgment on a trace. Feedback grades the agent's actual output. An expectation records the correct, ground-truth output. Labeling sessions capture either kind, and the docs say this data can then be used to improve your agent through systematic evaluation.

    Expectations do the most direct work. Databricks puts it this way: a trace that carries a ground-truth expectation becomes a ready-made test case for automated evaluation. You can promote it straight from the trace view, so the expert's correct answer becomes something every future version of the agent is tested against.

    Checkpoint 1 of 6· Put it in order

    Put these steps for turning a real interaction into an evaluation test case in order

    1. 1.Click Add to dataset to add the trace and its expectation to an evaluation dataset
    2. 2.Click Assess, then Add expectation to record the correct answer
    3. 3.In the Traces tab, click a trace to open it

    The labels live on the trace, so the developer doesn't have to do this alone. Anyone who reviews traces can add an expectation and promote it without leaving the trace UI. Feedback from end users of the deployed app works alongside this. It surfaces problematic queries and preserves successful interactions, and those traces are good candidates for the next round of expert review.

    Sources12

    2.Judge alignment: teaching LLM judges your experts' standards

    The second use of expert feedback is less obvious. It fixes the judges that score your agent. A generic LLM judge may disagree with your experts about what counts as relevant or correct. Judge alignment teaches LLM judges to match human evaluation standards through systematic feedback. According to the docs, this improves agreement with human assessments by 30 to 50 percent compared to baseline judges.

    Alignment works on built-in judges such as RelevanceToQuery, Safety, or Correctness, and on custom judges created with make_judge(). It needs MLflow 3.14.0 or above. There is one limit: alignment is not supported for session-level (multi-turn) judges such as ConversationCompleteness.

    Checkpoint 2 of 6· Put it in order

    Put the three steps of judge alignment in order

    1. 1.Collect human feedback: domain experts review and correct the judge's assessments
    2. 2.Generate initial assessments: run the judge on traces to establish a baseline
    3. 3.Align and deploy: call the judge's align() method to create a more aligned judge

    Checkpoint 3 of 6· Exam question

    A team's custom LLM judge for evaluating response correctness frequently disagrees with domain expert assessments on production traces. They want to systematically improve the judge's agreement with expert judgment without hand-tuning the judge's prompt. Which approach should they take?

    Sources3

    3.Logging expert feedback that alignment can use

    Alignment connects the expert's label to the judge's assessment by name. The human feedback assessment name must exactly match the judge's name attribute. For a built-in judge, that is the default snake_case name, such as relevance_to_query for RelevanceToQuery, unless you passed name= when creating it. For a custom judge, it is the name you gave make_judge(), such as product_quality. If an expert labels under any other name, alignment cannot match that label to the judge's assessment.

    Checkpoint 4 of 6· Check yourself

    You want to align a default RelevanceToQuery judge. Under which name must your experts log their feedback?

    There are two ways to collect the feedback. In Databricks UI review, experts open the experiment's Traces tab, review each trace and the judge's assessment, and add feedback under the matching name. Use programmatic feedback when you already have ground-truth labels, a large set of examples, or need collection to be reproducible.

    Choosing how to collect human feedback for alignment
    ApproachUse it when
    Databricks UI reviewYou need domain experts to review outputs, want to iteratively refine feedback criteria, or have a smaller dataset (< 100 examples)
    Programmatic feedbackYou have pre-existing ground truth labels, large datasets (100+ examples), or need reproducible feedback collection
    Logging existing ground-truth labels as HUMAN feedback under the judge's namepython
    # Log human feedback for each trace
    for item in ground_truth_data:
        mlflow.log_feedback(
            trace_id=item["trace_id"],
            name=initial_judge.name,  # Must match judge name (built-in or custom)
            value=item["label"],
            rationale=item.get("rationale", ""),
            source=AssessmentSource(
                source_type=AssessmentSourceType.HUMAN,
                source_id="ground_truth_dataset"
            ),
        )

    How much data, and of what kind? You can get reasonable alignment from at least 10 traces, but 50–100 traces give better results. The quality guidance is specific. Use diverse reviewers, meaning several domain experts. Keep the examples balanced, with at least 30% negative examples (poor or fair ratings). Ask for clear rationales that explain each rating. And choose representative samples that cover edge cases as well as common scenarios. A set made only of good answers teaches the judge almost nothing about what failure looks like.

    Sources3

    4.Running align() and registering the aligned judge

    Once enough traces carry both the judge's assessment and an expert's assessment, retrieve them and call align(). Built-in and custom judges use the same method. If you call align() without naming an optimizer, MLflow uses the MemAlign optimizer by default. Other optimizers are available in the package mlflow.genai.judges.optimizers.

    Checkpoint 5 of 6· Fill the gap

    Which method completes this alignment sample?

    if len(traces_for_alignment) >= 10:
        # Align the judge based on human feedback using the default optimizer
        aligned_judge = initial_judge. ? (traces_for_alignment)

    The aligned judge is a new object. Register it for production use under a new name so you can tell it apart from the original, and tag it with details such as how many traces it was aligned on. From then on, your evaluations and monitoring score the agent with a judge that reflects your experts' standards, not generic criteria.

    Registering the aligned judge under a distinct name with alignment metadatapython
        aligned_judge.register(
            experiment_id=experiment_id,
            name=f"{initial_judge.name}_aligned",
            tags={"alignment_date": "2025-10-23", "num_traces": str(len(traces_for_alignment))}
        )

    Checkpoint 6 of 6· Exam question

    An organization contracts an external subject matter expert to review and label agent traces in the Review App as part of a labeling session, but the expert should not receive broader access to the Databricks workspace. What is the minimum access the expert needs?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Any human feedback on a trace is used by align(), whatever name it was logged under.Why is that wrong?

      Alignment only pairs feedback whose assessment name exactly matches the judge's name attribute.

      Covered in Logging expert feedback that alignment can use

    2. 2.Every judge can be aligned, including multi-turn judges such as ConversationCompleteness.Why is that wrong?

      Alignment does not support session-level judges. It applies to built-in and custom judges that score individual traces.

      Covered in Judge alignment: teaching LLM judges your experts' standards

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “a trace that carries a ground-truth expectation becomes a ready-made test case for automated evaluation.”
      ↩︎ From expert expectations to evaluation datasets
      “Capture feedback from users of your deployed app to surface problematic queries and preserve successful interactions.”
      ↩︎ From expert expectations to evaluation datasets
      “Click Add to dataset to add the trace — with its expectation — to an evaluation dataset.”
      ↩︎ Checkpoint
    2. 2.
      “You can capture either Feedback or Expectation data, which can then be used to improve your agent through systematic evaluation.”
      ↩︎ From expert expectations to evaluation datasets
    3. 3.
      “Judge alignment teaches LLM judges to match human evaluation standards through systematic feedback.”
      ↩︎ Judge alignment: teaching LLM judges your experts' standards
      “improving agreement with human assessments by 30 to 50 percent compared to baseline judges.”
      ↩︎ Judge alignment: teaching LLM judges your experts' standards
      “The same alignment workflow applies to both built-in judges (such as RelevanceToQuery, Safety, or Correctness) and custom judges created with make_judge().”
      ↩︎ Judge alignment: teaching LLM judges your experts' standards
      “You can achieve reasonable alignment with at least 10 traces, but 50-100 traces yield better results.”
      ↩︎ Logging expert feedback that alignment can use
      “Balanced examples: Include at least 30% negative examples (poor/fair ratings)”
      ↩︎ Logging expert feedback that alignment can use
      “When you call align() without specifying an optimizer, the MemAlign optimizer is used automatically”
      ↩︎ Running align() and registering the aligned judge
      “The system supports the optimizers that are available in the package mlflow.genai.judges.optimizers.”
      ↩︎ Running align() and registering the aligned judge
      “The human feedback assessment name must exactly match the judge's name attribute.”
      ↩︎ Exam trap 1
      “Alignment is not supported for session-level (multi-turn) judges such as ConversationCompleteness.”
      ↩︎ Exam trap 2
      “Align and deploy: Invoke the judge's align() method to create a new judge that is more aligned with human feedback.”
      ↩︎ Prediction
      “Collect human feedback: Domain experts review and correct judge assessments.”
      ↩︎ Checkpoint
      “For built-in judges, this is the default snake_case name (for example, relevance_to_query for RelevanceToQuery)”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 6 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.