CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 55/56

    Custom Code-Based Scorers in MLflow: Signature, Inputs and Return Values

    Use Databricks custom Scorers for evaluating agents and LLMs

    9 min read
    1.79% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what an MLflow scorer receives, does and returns
    • Choose between built-in judges, custom LLM judges, code-based scorers and third-party scorers
    • Write a @scorer function that declares only the inputs it needs
    • Predict which return type and metric name a scorer produces in the results

    Key concept

    Scorer — A scorer is an evaluation function. It takes an MLflow trace of one app interaction, pulls out the fields it needs, judges quality, and attaches the result to that trace as Feedback. A code-based custom scorer is a scorer whose logic you write yourself in Python.

    1.What a scorer does with a trace

    In the MLflow GenAI evaluation framework, scorers are the single interface for defining evaluation criteria for models, agents and applications. Every scorer follows the same three steps. It receives a Trace, from either mlflow.genai.evaluate() during development or the monitoring service in production. It parses the trace to extract the fields it needs. It runs its assessment and returns the result as Feedback attached to the trace.

    The result does not have to be a number. It could be a pass/fail, true/false, numerical value, or a categorical value. Because scorers work on traces and not on a particular dataset format, you can use the same scorer to evaluate in development and to monitor in production. That keeps your evaluation consistent across the application lifecycle.

    Checkpoint 1 of 6· Check yourself

    After a scorer parses a trace and runs its quality assessment, what does it do with the result?

    Sources1

    2.Where custom scorers sit among the scorer types

    Databricks describes scorer types as a ladder. Each step adds customization and complexity. The recommended path is to start with built-in judges, move to custom LLM judges for domain-specific criteria, and write code-based scorers for deterministic business logic.

    Scorer approaches, from least to most customizable
    ApproachLevel of customizationTypical use
    Built-in judgesMinimal (moderate for Guidelines judges)Quick LLM evaluation with judges such as Correctness and RetrievalGroundedness
    Custom judgesFullFully customized LLM judges that can return numerical scores, categories or boolean values
    Code-based scorersFullProgrammatic, deterministic checks: exact matching, format validation, performance metrics
    Third-party scorersFullSpecialized metrics from open-source evaluation frameworks

    Custom LLM judges let you write evaluation criteria in natural language. Code-based scorers are for cases where neither built-in nor custom LLM judges fit. The documentation names four reasons to write one:

    - to define a custom heuristic or programmatic metric - to change how trace data is mapped to a built-in LLM judge - to use your own LLM in place of a Databricks-hosted judge - to get more flexibility and control than a custom LLM judge offers

    Checkpoint 2 of 6· Check yourself

    A team needs to verify that every agent response is valid JSON with three required fields, and the result must be deterministic. Which approach fits best?

    Sources21

    3.The @scorer signature and its four inputs

    Most code-based scorers should be written as ordinary Python functions with the @scorer decorator. MLflow passes the complete trace, and it also pulls out the commonly needed parts and passes them as keyword-only named arguments:

    Full @scorer signature: every argument is optional and keyword-onlypython
    from mlflow.genai.scorers import scorer
    from typing import Optional, Any
    from mlflow.entities import Feedback
    
    @scorer
    def my_custom_scorer(
        *,  # All arguments are keyword-only
        inputs: Optional[dict[str, Any]],       # App's raw input, a dictionary of input argument names and values
        outputs: Optional[Any],                 # App's raw output
        expectations: Optional[dict[str, Any]], # Ground truth, a dictionary of label names and values
        trace: Optional[mlflow.entities.Trace]  # Complete trace with all spans and metadata
    ) -> Union[int, float, bool, str, Feedback, List[Feedback]]:
        # Your evaluation logic here

    The four inputs are:

    - inputs: the request sent to the app - outputs: the app's response - expectations: ground truth or labels - trace: the full trace, passed as an instantiated mlflow.entities.trace class, so you can analyse intermediate steps, latency and tool usage

    When you run mlflow.genai.evaluate(), the first three can come from the data argument or be parsed from the trace. You only declare the arguments your logic uses. A word-count metric needs nothing but outputs.

    Checkpoint 3 of 6· Check yourself

    You write @scorer def response_length(outputs: str) -> int:. Which statement is correct?

    Sources3

    4.Return types and how metrics get their names

    For simple pass/fail or numeric results, return a primitive value. The MLflow UI shows each type differently:

    Two simple scorers: one returns a number, one returns a yes/no pass/failpython
    @scorer
    def response_length(outputs: str) -> int:
        # Return a numeric metric
        return len(outputs.split())
    
    @scorer
    def contains_citation(outputs: str) -> str:
        # Return pass/fail string
        return "yes" if "[source]" in outputs else "no"

    For a detailed assessment, return a Feedback object. It carries a value plus an optional rationale, source (an AssessmentSource such as HUMAN, CODE or LLM_JUDGE) and metadata. To score several aspects in one pass, return a list of Feedback objects. Each one should set name, and each name appears as a separate metric.

    Checkpoint 4 of 6· Match them up

    Match each scorer return type to how the MLflow UI shows it

    Tap a term, then the definition that fits it.

    Names matter because they become metric names in evaluation results, monitoring results and dashboards. Use clear names such as safety_check or relevance_monitor. The rules are:

    - A primitive value or an unnamed Feedback takes the function name with @scorer, or the Scorer.name field with the class. - A named Feedback always uses its own name. - All metrics must have distinct names, and that includes each Feedback in a returned list.

    Checkpoint 5 of 6· Check yourself

    A @scorer function named quality_gate returns Feedback(value=True) with no name. What is the metric called?

    Checkpoint 6 of 6· Exam question

    An evaluation team needs a check that confirms each agent response contains a syntactically valid JSON object matching a required schema — a deterministic pass/fail check that does not need any LLM judgment. Which approach fits this requirement, and what should it return?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The metric name of a @scorer is always the Python function name.Why is that wrong?

      If the scorer returns Feedback objects with name set, those names win. The function name is used only for primitive values or unnamed Feedback.

      Covered in Return types and how metrics get their names

    2. 2.Custom code-based scorers can only compute heuristics and cannot involve an LLM.Why is that wrong?

      One documented use for a code-based scorer is calling your own LLM in place of a Databricks-hosted judge. Another is changing how trace data is mapped to a built-in judge.

      Covered in Where custom scorers sit among the scorer types

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “A scorer receives a Trace from either evaluate() or the monitoring service.”
      ↩︎ What a scorer does with a trace
      “This could be a pass/fail, true/false, numerical value, or a categorical value.”
      ↩︎ What a scorer does with a trace
      “create custom code-based scorers for deterministic business logic”
      ↩︎ Where custom scorers sit among the scorer types
      “Returns the quality assessment as Feedback to attach to the trace”
      ↩︎ Key concept
      “Using your own LLM (rather than a Databricks-hosted LLM judge) for evaluation.”
      ↩︎ Exam trap 2
      “Returns the quality assessment as Feedback to attach to the trace”
      ↩︎ Checkpoint
      “Programmatic and deterministic scorers that evaluate things like exact matching, format validation, and performance metrics.”
      ↩︎ Checkpoint
    2. 2.
      “Use them when built-in LLM judges and custom LLM judges don't fit your evaluation needs.”
      ↩︎ Where custom scorers sit among the scorer types
    3. 3.
      “Scorers receive the complete MLflow trace containing all spans, attributes, and outputs.”
      ↩︎ The @scorer signature and its four inputs
      “The trace is passed to the custom scorer as an instantiated mlflow.entities.trace class.”
      ↩︎ The @scorer signature and its four inputs
      “Each feedback should have the name field specified, and those names are displayed as separate metrics in the evaluation results.”
      ↩︎ Return types and how metrics get their names
      “For evaluation and monitoring, all metrics must have distinct names.”
      ↩︎ Return types and how metrics get their names
      “If the scorer returns one or more Feedback objects, then Feedback.name fields take precedence, if specified.”
      ↩︎ Exam trap 1
      “All input arguments are optional, so declare only what your scorer needs”
      ↩︎ Checkpoint
      “If the scorer returns one or more Feedback objects, then Feedback.name fields take precedence, if specified.”
      ↩︎ Prediction
      “Return a Feedback object or list of Feedback objects for detailed assessments with scores, rationales, and metadata.”
      ↩︎ Checkpoint
      “For primitive return values or unnamed Feedbacks, the function name (for the @scorer decorator) or the Scorer.name field (for the Scorer class) is used.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Building and Running Custom Scorers: Trace Patterns, Errors and Production Monitoring

    Spotted a mistake, or was something unclear? Tell us.