CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 49/56

    MLflow Scorers and LLM Judges for Agent Evaluation

    Evaluate agent performance using MLflow scoring and tracing

    10 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Choose between built-in judges, custom judges, code-based scorers and third-party scorers
    • Identify which built-in judges need ground truth (expectations) and which do not
    • Write Guidelines judges and change the model a judge uses
    • Use evaluation results to improve an agent and check that nothing regressed

    1.Four kinds of scorer, from least to most control

    A scorer returns a judgement on some aspect of quality. The judgement can be pass/fail, true/false, a number or a category. Databricks presents the scorer options as a ladder. Start with built-in judges, add custom LLM judges for criteria specific to your domain, and write code-based scorers for deterministic business logic. LLM judges are scorers that use an LLM to make the assessment. They can look at inputs, outputs and the whole execution trace, and they recognise when two differently worded queries mean the same thing. The same scorer object works in development evaluation and in production monitoring, which keeps quality measured the same way across the lifecycle.

    Scorer approaches compared
    ApproachCustomizationTypical use
    Built-in judgesMinimal (moderate for Guidelines judges)Fast LLM evaluation, e.g. Correctness, RetrievalGroundedness
    Custom judgesFullDomain-specific LLM criteria returning numbers, categories or booleans
    Code-based scorersFullDeterministic checks such as exact matching, format validation, performance metrics
    Third-party scorersFullSpecialized metrics from open-source evaluation frameworks

    Checkpoint 1 of 3· Check yourself

    Every agent response must be valid JSON with a fixed set of keys. Which scorer approach fits best?

    Sources1

    2.Built-in judges and the ground-truth question

    Single-turn built-in judges: which ones need expectations (ground truth)
    JudgeRequires ground truthWhat it evaluates
    RelevanceToQueryNoResponse is relevant to the user's request
    RetrievalRelevanceNoRetrieved context is relevant to the request
    SafetyNoContent is free of harmful or toxic material
    RetrievalGroundednessNoResponse is grounded in retrieved context (hallucination)
    GuidelinesNoResponse meets specified natural-language criteria
    ToolCallEfficiencyNoTool calls are free of redundancy
    ExpectationsGuidelinesNo (but needs guidelines in expectations)Per-example natural-language criteria
    CorrectnessYesResponse is correct compared with ground truth
    RetrievalSufficiencyYesContext contains everything needed for the ground-truth facts
    ToolCallCorrectnessYesTool calls and arguments are correct for the query

    The pattern is that a judge needs ground truth when it compares against a known right answer: Correctness, RetrievalSufficiency and ToolCallCorrectness. Reference-free judges can therefore score unlabeled traffic, while ground-truth judges need an evaluation dataset that has expectations filled in. This matters for production monitoring: registered monitoring scorers always parse inputs and outputs from the trace, and expectations is not available to them. So the ground-truth judges cannot get their labels in monitoring, and reference-free judges are the ones that fit live traffic. For conversational agents there are also multi-turn judges that take a whole session, such as ConversationCompleteness, UserFrustration and KnowledgeRetention. None of these needs ground truth.

    The retrieval judges also depend on how you instrument the app. They need to find the retriever step in the trace, so the Databricks tutorial marks its retrieval function with span_type="RETRIEVER".

    Checkpoint 2 of 3· Match them up

    Match each situation to the built-in judge that fits it

    Tap a term, then the definition that fits it.

    Sources234

    3.Guidelines, judge models and code-based scorers

    Guidelines judges are the quickest way to encode your own rules. Each one is a named, natural-language rule that the judge marks pass or fail. In the Databricks email-agent tutorial, RetrievalGroundedness, RelevanceToQuery and Safety run alongside five Guidelines judges, for instruction following, conciseness, naming the contact, professional tone and concrete next steps. All of them are passed together as scorers to mlflow.genai.evaluate().

    One Guidelines judge from the tutorial's judge listpython
    Guidelines(
                name="professional_tone",
                guidelines="The email must be in a professional tone.",
            ),

    By default, each judge runs on a Databricks-hosted LLM designed for quality assessment. To use a different model, pass the model argument in the form <provider>:/<model-name>.

    Changing the LLM that powers a built-in judgepython
    from mlflow.genai.scorers import Correctness
    
    Correctness(model="databricks:/databricks-gpt-5-mini")

    When neither of those is enough, there are two further options. Custom LLM judges give you control over grades and scores beyond pass/fail. They can also check whether the agent made the right decisions, and judge alignment can tune them to match human evaluation standards. Custom code-based scorers are defined with the @scorer decorator or the Scorer class. Use them for custom heuristics, for changing how trace data is mapped to a built-in judge, or for running evaluation on your own LLM instead of a Databricks-hosted judge.

    Most code-based scorers use the @scorer decorator. A scorer receives the complete trace, and MLflow also passes commonly needed data as named, keyword-only arguments: inputs (the request sent to the app), outputs (the app's response), expectations (ground truth or labels) and trace (all spans, for analysing intermediate steps, latency or tool use). All of them are optional, so declare only what your scorer needs. In mlflow.genai.evaluate(), inputs, outputs and expectations can come from the data argument or be parsed from the trace.

    A code-based scorer that returns a "yes"/"no" stringpython
    @scorer
    def contains_citation(outputs: str) -> str:
        # Return pass/fail string
        return "yes" if "[source]" in outputs else "no"
    What a code-based scorer can return
    Return TypeMLflow UI DisplayUse Case
    "yes"/"no"Pass/FailBinary evaluation
    True/FalseTrue/FalseBoolean checks
    int/floatNumeric valueScores, counts
    FeedbackValue + rationaleDetailed assessment
    List[Feedback]Multiple metricsMulti-aspect evaluation

    To run a scorer on live traffic, register it with the experiment and then start it with a sampling configuration. This .register() then .start() pattern works for built-in judges, custom judges, code-based scorers and multi-turn judges. sample_rate sets the fraction of traces the scorer evaluates. Databricks recommends sample_rate=1.0 for critical scorers such as safety and security checks, and lower rates (0.05-0.2) for expensive scorers such as complex LLM judges. Two limits apply. Registered monitoring scorers parse inputs and outputs from the trace and have no expectations. Only @scorer-decorated functions, defined and registered from a Databricks notebook, are supported: class-based Scorer subclasses are not.

    Registering and starting a built-in judge for production monitoringpython
    from mlflow.genai.scorers import Safety, ScorerSamplingConfig
    
    # Register and start a built-in judge
    safety_judge = Safety().register(name="safety")
    safety_judge = safety_judge.start(sampling_config=ScorerSamplingConfig(sample_rate=0.7))

    Sources135

    4.From scores to a better agent

    Scores only matter if they lead to changes. In the tutorial's first run, the feedback shows three failure patterns: weak instruction following, emails that are too long, and missing concrete next steps. The fixes target those patterns directly. Options include prompt engineering, guardrails that check outputs before users see them, retrieval improvements found by examining retrieval spans, and splitting reasoning into more spans. The judge list is saved in a variable, so the improved version is evaluated with exactly the same scorers. That makes the before-and-after comparison fair.

    Checkpoint 3 of 3· Put it in order

    Put the tutorial's evaluate-and-improve workflow in order

    1. 1.Interpret results to identify quality issues
    2. 2.Evaluate quality with LLM judges using the evaluation harness
    3. 3.Improve the app based on evaluation results
    4. 4.Compare versions to verify improvements and catch regressions
    5. 5.Create evaluation datasets from real usage data

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Hallucination checks such as RetrievalGroundedness need a labeled expected answer.Why is that wrong?

      RetrievalGroundedness compares the response with the retrieved context and needs only inputs and outputs. It is Correctness that compares against ground truth.

      Covered in Built-in judges and the ground-truth question

    2. 2.Any built-in judge can run on unlabeled data.Why is that wrong?

      Correctness, RetrievalSufficiency and ToolCallCorrectness all require ground truth in expectations, so the dataset must provide it.

      Covered in Built-in judges and the ground-truth question

    3. 3.A scorer registered for production monitoring can read expectations, so Correctness works on live traffic.Why is that wrong?

      Registered monitoring scorers parse inputs and outputs from the trace and have no expectations, so ground-truth judges cannot get their labels there.

      Covered in Guidelines, judge models and code-based scorers

    Practise it for real

    Score pre-computed agent answers with built-in judges, then inspect the feedback on the resulting traces

    1. 1.Install mlflow[databricks]>=3.1.0 and set up an MLflow experiment

      Why: Evaluation results are stored as traces in the active experiment

      You should see: An experiment you can open under Experiments

    2. 2.Build a list of dicts, each with an inputs question and a pre-computed outputs response

      Why: This is answer sheet evaluation, so no predict_fn is needed

      You should see: Two or more rows with inputs and outputs keys

    3. 3.Call mlflow.genai.evaluate(data=..., scorers=[Safety(), RelevanceToQuery()])

      Why: Both judges are reference-free, so the data needs no expectations

      You should see: An EvaluationResult with a run_id

    4. 4.Call mlflow.search_traces(run_id=<result>.run_id) and print the assessments column

      Why: Judge feedback is attached to each trace

      You should see: One trace per row, each with Safety and RelevanceToQuery feedback

    Stuck? Get a nudge

    Try adding Correctness() without expectations, and notice that it needs ground truth that your rows do not contain.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “You can use the same scorer for evaluation in development and monitoring in production to keep evaluation consistent throughout the application lifecycle.”
      ↩︎ Four kinds of scorer, from least to most control
      “Specify the model in the format <provider>:/<model-name>.”
      ↩︎ Guidelines, judge models and code-based scorers
      “Customizing how the data from your app's trace is mapped to built-in LLM judges.”
      ↩︎ Guidelines, judge models and code-based scorers
      “Programmatic and deterministic scorers that evaluate things like exact matching, format validation, and performance metrics.”
      ↩︎ Checkpoint
    2. 2.
      “Is the response correct as compared to the provided ground truth?”
      ↩︎ Built-in judges and the ground-truth question
      “Correctness | inputs, outputs, expectations | Yes”
      ↩︎ Exam trap 1
      “Does the context provide all necessary information to generate a response that includes the ground truth facts?”
      ↩︎ Exam trap 2
    3. 3.
      “Registered scorers for production monitoring always parse the inputs and outputs parameters from the trace. expectations is not available.”
      ↩︎ Built-in judges and the ground-truth question
      “Define custom code-based scorers in MLflow using the @scorer decorator or the Scorer class.”
      ↩︎ Guidelines, judge models and code-based scorers
      “All input arguments are optional, so declare only what your scorer needs:”
      ↩︎ Guidelines, judge models and code-based scorers
      “Registered scorers for production monitoring always parse the inputs and outputs parameters from the trace. expectations is not available.”
      ↩︎ Exam trap 3
    4. 4.
      “The retrieval component is marked with span_type="RETRIEVER" to enable MLflow's retrieval-specific LLM judges.”
      ↩︎ Built-in judges and the ground-truth question
      “Retrieval improvements (for RAG apps): Enhance retrieval mechanisms if relevant documents aren't being found by examining retrieval spans”
      ↩︎ From scores to a better agent
      “Compare versions to verify improvements worked and did not cause regressions.”
      ↩︎ Checkpoint
    5. 5.
      “For critical scorers such as safety and security checks, use sample_rate=1.0.”
      ↩︎ Guidelines, judge models and code-based scorers

    Ready to test yourself?

    Practise the 5 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.