CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 6 · Lesson 55/56

    Building and Running Custom Scorers: Trace Patterns, Errors and Production Monitoring

    Use Databricks custom Scorers for evaluating agents and LLMs

    12 min read
    1.79% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Iterate on a scorer against stored traces without rerunning the app
    • Write scorers that inspect trace spans or wrap a built-in LLM judge
    • Handle missing expectations and scorer errors correctly
    • Know which scorer definitions production monitoring supports, and register one that reads secrets

    1.Iterating on a scorer against stored traces

    A code-based scorer is a Python function decorated with @scorer that receives an MLflow trace and returns a value or Feedback. To get it right you usually need several attempts, and calling an LLM-backed app on every attempt is slow. Databricks recommends a workflow that runs the app once and reuses its traces:

    1. Define evaluation data. 2. Generate traces from your app. 3. Query and store the resulting traces. 4. Iterate on the scorer against the stored traces.

    After that run, mlflow.search_traces(run_id=eval_results.run_id) returns the traces as a Pandas DataFrame. You pass that DataFrame straight back to evaluate() as data and leave out predict_fn. Only the new scorer runs, and the app is not called.

    Scoring precomputed traces: no predict_fn, so the app is not called againpython
    from mlflow.genai.scorers import scorer
    
    @scorer
    def response_length(outputs: str) -> int:
        # Example metric.
        # Implement your actual metric logic here.
        return len(outputs)
    
    # Note the lack of a predict_fn parameter.
    mlflow.genai.evaluate(
        data=generated_traces,
        scorers=[response_length],
    )

    Checkpoint 1 of 8· Put it in order

    Put the scorer development workflow in order

    1. 1.Run evaluate() on the stored traces with the scorer you are iterating on
    2. 2.Generate traces by running evaluate() with predict_fn and a placeholder scorer
    3. 3.Define evaluation data
    4. 4.Store the traces with mlflow.search_traces()

    Sources1

    2.Two common patterns: reading spans and wrapping a judge

    When the scorer declares the trace argument, it can read individual spans with their inputs, outputs, attributes and timing. The documented example finds the chat-model span and checks that the LLM call finished within 5 seconds. It returns "yes" or "no" as Feedback, with a rationale that states the measured time.

    Checkpoint 2 of 8· Fill the gap

    Which Trace method finds the chat-model span in this latency scorer?

    @scorer
    def llm_response_time_good(trace: Trace) -> Feedback:
        # Search particular span type from the trace
        llm_span = trace. ? (span_type=SpanType.CHAT_MODEL)[0]
    
        response_time = (llm_span.end_time_ns - llm_span.start_time_ns) / 1e9 # convert to seconds

    The second pattern wraps a built-in LLM judge. Use it to preprocess trace data for the judge or to post-process the judge's feedback. In the example, the app's inputs is a whole chat history: {"messages": [...]}. The relevance judge needs a single request, so the scorer walks the messages backwards to find the last user message. It then calls is_context_relevant and returns that judge's Feedback unchanged.

    End of the wrapper: it fails loudly if no user message exists, otherwise it hands the extracted request to the built-in judgepython
        if not last_user_message_content:
            raise Exception("Could not extract the last user message from inputs to evaluate relevance.")
    
        # Call the `relevance_to_query judge. It will return a Feedback object.
        return is_context_relevant(
            request=last_user_message_content,
            context={"response": outputs},
        )

    Sources2

    3.Expectations: available offline, missing in production

    Expectations are ground-truth values or labels. When you run mlflow.genai.evaluate(), you can supply them in two ways. One is an expectations key on each row of a list or DataFrame. The other is a trace field, as in the DataFrame from mlflow.search_traces(), which carries any Expectation data attached to the traces.

    Production monitoring works differently. Registered scorers always parse inputs and outputs from the trace, and expectations is not available, because live traffic has no ground truth. A scorer that you reuse in both places must not assume that expectations is filled in.

    Checkpoint 3 of 8· Check yourself

    A scorer compares outputs with expectations["expected_response"]. It passes offline evaluation, and you then register it for production monitoring. What should you expect?

    Checkpoint 4 of 8· Exam question

    A team writes a custom scorer for a support agent and wants stakeholders reviewing evaluation runs to see not just a numeric score but also a short written reason for that score, attached to each trace. What should the `@scorer`-decorated function return to achieve this?

    Sources2

    4.When a scorer throws

    If a scorer fails on one trace, MLflow records the error for that trace and carries on with the others. The recommended approach is to let exceptions propagate. MLflow catches the exception and creates a Feedback whose value is None and whose error holds the exception details, and you can open that row in the results to see them. When you want your own message, catch the exception and return a Feedback with error set. That can be the exception object itself, or an AssessmentError with a structured error_code.

    Explicit error handling: a structured AssessmentError for missing fields, and the raw exception for invalid JSONpython
            if missing:
                return Feedback(
                    error=AssessmentError(
                        error_code="MISSING_REQUIRED_FIELDS",
                        error_message=f"Missing required fields: {missing}",
                    ),
                )
    
            return Feedback(
                value=True,
                rationale="Valid JSON with all required fields"
            )
    
        except json.JSONDecodeError as e:
            return Feedback(error=e)  # Can pass exception object directly to the error parameter

    Checkpoint 5 of 8· Check yourself

    A custom scorer raises an uncaught KeyError on 3 of 200 traces. What does MLflow do?

    Sources3

    5.@scorer vs the Scorer class, and running scorers in production

    Two ways to define a code-based scorer
    ApproachUse whenProduction monitoring
    @scorer decoratorMost cases. Recommended starting point.Supported (when defined and registered from a Databricks notebook)
    Scorer classYou need stateful scorers, complex initialization, or Pydantic fieldsNot supported

    The Scorer class is a Pydantic object. You must set its name field, add any extra fields you need, and put your logic in __call__. Keep state in instance attributes, not mutable class attributes: a class-level results = [] is shared by every instance.

    Correct per-instance state in a Scorer subclasspython
    # CORRECT: Use instance attributes
    class GoodScorer(Scorer):
        results: list[str] = None
    
        name: str = "good_scorer"
    
        def __init__(self):
            self.results = []  # Per-instance state
    
        def __call__(self, outputs, **kwargs):
            self.results.append(outputs)  # Safe
            return Feedback(value=True)

    Production monitoring accepts built-in LLM judges and @scorer functions. It does not accept Scorer subclasses. If you need state in production, keep it inside the body of a decorated function. The function must also be defined and registered from a Databricks notebook, because the monitoring service serializes its code to run it remotely.

    Scorers that call an external LLM endpoint can read credentials from Databricks secrets, both in development and in production. dbutils is not available in the scorer runtime by default, so import it inside the function.

    Checkpoint 6 of 8· Fill the gap

    Which module must the scorer import dbutils from to read secrets?

    @scorer
    def custom_llm_scorer(trace: Trace) -> Feedback:
        # Explicitly import dbutils to access secrets
        from  ?  import dbutils
    
        # Retrieve your API key from Databricks secrets
        api_key = dbutils.secrets.get(scope='my-scope', key='api-key')

    To deploy a decorated scorer, register it and then start it with a sampling configuration. A sample_rate of 1 scores every trace.

    Registering and starting a scorer for production monitoringpython
    # Register and start the scorer
    custom_llm_scorer.register()
    custom_llm_scorer.start(sampling_config = ScorerSamplingConfig(sample_rate=1))

    Checkpoint 7 of 8· Check yourself

    A team built a stateful scorer as a Scorer subclass and now wants to run it on live traffic. What should they do?

    Checkpoint 8 of 8· Exam question

    A team is writing a custom scorer that must flag agent runs where a particular tool span took longer than 2 seconds, using the full execution trace rather than just the final input and output text. Which parameter should the scorer function declare to access this span-level timing information?

    Sources43

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Any custom scorer that works in mlflow.genai.evaluate() can be registered for production monitoring.Why is that wrong?

      Production monitoring supports built-in judges and @scorer functions registered from a Databricks notebook. Scorer subclasses are not supported.

      Covered in @scorer vs the Scorer class, and running scorers in production

    2. 2.A scorer that depends on expectations works the same way in production monitoring as in offline evaluation.Why is that wrong?

      Live traffic has no ground truth. Registered production scorers receive inputs and outputs only, so a reused scorer has to cope without expectations.

      Covered in Expectations: available offline, missing in production

    3. 3.dbutils is available inside a scorer just as it is in a notebook cell, so secrets can be read directly.Why is that wrong?

      The scorer runtime has no dbutils by default. You must import it from databricks.sdk.runtime inside the scorer function.

      Covered in @scorer vs the Scorer class, and running scorers in production

    Practise it for real

    Develop a custom scorer against stored traces without rerunning your app on every change

    1. 1.Create an eval_dataset list whose rows each have an "inputs" dict, and a traced app function (decorated with @mlflow.trace)

      Why: Evaluation needs input rows, and tracing records what the scorer will read

      You should see: A list of rows plus a callable app

    2. 2.Run mlflow.genai.evaluate(data=eval_dataset, predict_fn=sample_app, scorers=[placeholder_metric]), where placeholder_metric is a @scorer that returns 1

      Why: evaluate() needs at least one scorer, and this run exists only to produce traces

      You should see: One trace in the experiment for each dataset row

    3. 3.Store the traces with generated_traces = mlflow.search_traces(run_id=eval_results.run_id)

      Why: The stored traces become reusable evaluation input

      You should see: A Pandas DataFrame of traces

    4. 4.Write a real @scorer, for example response_length on outputs, and run mlflow.genai.evaluate(data=generated_traces, scorers=[response_length]) without predict_fn

      Why: Leaving out predict_fn scores the stored traces without calling the app

      You should see: A response_length metric for each trace, with no new LLM calls from the app

    Stuck? Get a nudge

    If outputs is not a plain string for your app, declare trace instead and read the span you need.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Use this developer workflow to update your scorer without rerunning your entire app each time”
      ↩︎ Iterating on a scorer against stored traces
      “The mlflow.search_traces() function returns a Pandas DataFrame of traces.”
      ↩︎ Iterating on a scorer against stored traces
      “Since evaluate() requires at least one scorer, define a placeholder scorer for this initial trace generation”
      ↩︎ Prediction
      “This allows you to quickly iterate on your metric without having to re-run your app.”
      ↩︎ Checkpoint
    2. 2.
      “Access the full MLflow Trace object to use various details (spans, inputs, outputs, attributes, timing) for fine-grained metric calculation.”
      ↩︎ Two common patterns: reading spans and wrapping a judge
      “Use this to preprocess trace data for the judge or post-process its feedback.”
      ↩︎ Two common patterns: reading spans and wrapping a judge
      “design it to handle expectations gracefully”
      ↩︎ Expectations: available offline, missing in production
      “Production monitoring typically doesn't have expectations since you are evaluating live traffic without ground truth.”
      ↩︎ Exam trap 2
    3. 3.
      “MLflow automatically captures the exception and creates a Feedback object with the following error details”
      ↩︎ When a scorer throws
      “AssessmentError: For structured error reporting with error codes.”
      ↩︎ When a scorer throws
      “be sure to use instance attributes, not mutable class attributes.”
      ↩︎ @scorer vs the Scorer class, and running scorers in production
      “This approach works for both development evaluation and production monitoring.”
      ↩︎ @scorer vs the Scorer class, and running scorers in production
      “By default, dbutils isn't available in the scorer runtime environment.”
      ↩︎ Exam trap 3
      “Registered scorers for production monitoring always parse the inputs and outputs parameters from the trace. expectations is not available.”
      ↩︎ Checkpoint
      “Let exceptions propagate (recommended) so that MLflow can capture error messages for you.”
      ↩︎ Checkpoint
    4. 4.
      “@scorer-decorated functions used in production monitoring must be defined and registered from a Databricks notebook.”
      ↩︎ @scorer vs the Scorer class, and running scorers in production
      “Class-based Scorer subclasses are not supported for production monitoring.”
      ↩︎ Exam trap 1
      “If you need stateful scorers in production, use the @scorer decorator and manage state inside the function body.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 6 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.