CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 26/56

    GenAI Evaluation Phase: Scoring an Agent Before Release

    Compare the evaluation and monitoring phases of the Gen AI application life cycle

    9 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain where the evaluation phase sits in the Gen AI application life cycle and how it differs from monitoring in timing
    • Describe what scorers and LLM judges do and why the same scorer can serve both evaluation and monitoring
    • Choose between running mlflow.genai.evaluate() with a predict_fn and scoring pre-computed outputs or existing traces
    • Identify when evaluation runs: version comparisons, prompt or model changes, and pre-release regression checks

    Key concept

    Scorer (shared across evaluation and monitoring) — A scorer is an LLM judge or code-based check that reads a trace and attaches a quality assessment to it as feedback. The same scorer runs during development evaluation and production monitoring, so the two phases use the same definition of quality.

    1.Two phases of one quality loop

    The exam asks you to tell apart two phases of quality work on a Gen AI application. The Databricks RAG cookbook separates them by time: evaluation happens during development, and monitoring happens once the application is deployed to production. The same sentence adds that the fundamental components are similar. Both phases need a definition of quality, a way to score responses against it, and traces that show how each response was produced. The differences are when the scoring runs, what data it runs over, and what you do with the result.

    MLflow treats the two phases as stages of one continuous loop, not as a hand-off from one team to another. You trace the agent and find an issue. You collect feedback and curate a dataset, then write a scorer that catches the issue automatically. You fix the agent and re-run evaluation to confirm the fix. Then you monitor production with the same scorers so the issue doesn't come back. Monitoring then surfaces the next issue, and the loop starts again. Evaluation sits in the middle of the loop and monitoring at its end. The scorer is what carries over from one to the other.

    Checkpoint 1 of 4· Put it in order

    Put these stages of the MLflow agent-improvement loop in order.

    1. 1.Write a scorer that catches the issue
    2. 2.Evaluate the fix by re-running evaluation against the dataset
    3. 3.Collect feedback and curate an evaluation dataset from failing traces
    4. 4.Fix the agent's prompt or logic
    5. 5.Monitor production by running the scorers against live traffic

    Sources12

    2.What evaluation measures: scorers and judges

    Evaluation answers one question: is the agent actually good? In the documentation's words, is the response correct, grounded, safe, and complete? MLflow answers it with scorers, which are LLM judges and code-based checks that attach a quality assessment to each trace. A scorer can return pass/fail, true/false, a number, or a category.

    LLM judges are central here because Gen AI output is open-ended. The RAG cookbook points out that reading every response on every evaluation run isn't feasible. A separate LLM reviewing the outputs scales the work and can check things human raters struggle with at volume, such as whether a response is grounded in thousands of tokens of retrieved context. MLflow lets you start simple and add control as your needs grow:

    Scorer approaches in MLflow GenAI evaluation, from least to most customization
    ApproachCustomizationTypical use
    Built-in judgesMinimal (moderate for Guidelines judges)Quick LLM evaluation with scorers such as Correctness and RetrievalGroundedness; Guidelines judges check pass/fail natural-language rules
    Custom judgesFullDomain-specific LLM judges returning numerical scores, categories, or boolean values
    Code-based scorersFullDeterministic checks such as exact matching, format validation, and performance metrics
    Third-party scorersFullSpecialized metrics from open-source evaluation frameworks

    The way a scorer works is what links the two phases. It receives a trace, parses out the fields it needs, runs its assessment, and returns the result as Feedback that is attached to the trace. The scorer doesn't care whether the trace came from an evaluation run or from the monitoring service. That's why one scorer can serve both phases.

    Checkpoint 2 of 4· Check yourself

    A team has written a custom code-based scorer during development. Where can that scorer receive traces from?

    Sources314

    3.Running the evaluation phase with mlflow.genai.evaluate()

    In development, evaluation runs against an evaluation set: a curated set of queries (ideally with expected outputs) that represents how the application will be used. The cookbook says these examples should be challenging, diverse, and updated as usage changes. An evaluation harness then runs the application on every record and passes each output through the judges. In MLflow that harness is mlflow.genai.evaluate(). You can start by scoring a few traces from the UI, then move the same scorers into a repeatable evaluate() loop as the app matures.

    The mlflow.genai.evaluate() signature: test data, scorers, an optional predict_fn, and an optional model_id for version trackingpython
    def mlflow.genai.evaluate(
        data: Union[pd.DataFrame, List[Dict], mlflow.genai.datasets.EvaluationDataset],  # Test data.
        scorers: list[mlflow.genai.scorers.Scorer],  # Quality metrics, built-in or custom.
        predict_fn: Optional[Callable[..., Any]] = None,  # App wrapper. Used for direct evaluation only.
        model_id: Optional[str] = None,  # Optional version tracking.
    ) -> mlflow.models.evaluation.base.EvaluationResult:

    The documentation lists three typical occasions for evaluation: nightly or weekly checks against curated datasets, validating prompt or model changes across app versions, and checks before a release or PR to prevent quality regressions. All three are deliberate, on-demand runs that compare one version against another. Evaluation is how you choose a version before users see it.

    Beyond calling a predict_fn, evaluate() has an "answer sheet" mode for cases where you can't, or don't want to, run the agent during evaluation. You supply inputs and pre-computed outputs, or existing traces, and evaluate() runs the scorers and logs an evaluation run. The example below pulls traces from production and scores them in a batch. This is still evaluation: it is a run you trigger, over a fixed set of traces, and it produces an evaluation result.

    Evaluating existing production traces in a one-off batch with mlflow.genai.evaluate()python
    import mlflow
    
    # Retrieve traces from production
    traces = mlflow.search_traces(
        filter_string="trace.status = 'OK'",
    )
    
    # Evaluate problematic traces
    evaluation = mlflow.genai.evaluate(
        data=traces,
        scorers=[Safety(), RelevanceToQuery()]
    )

    Answer-sheet mode has one catch. If the answer sheet contains traces that differ from what your production environment emits, a scorer that parses them may not work on live traces. The documentation warns you may need to rewrite your scorer functions before using them for production monitoring. Using your production app as the predict_fn avoids that mismatch.

    Checkpoint 3 of 4· Exam question

    A team building a customer support RAG agent on Databricks has just finished revising a prompt template. Before shipping the change to production, they want to systematically compare the new version's response quality against a curated set of question-answer pairs and flag regressions relative to the prior version. Which approach best fits this need?

    Checkpoint 4 of 4· Check yourself

    Which situation is a typical use of mlflow.genai.evaluate() during the evaluation phase?

    Sources135

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Using mlflow.genai.evaluate() on traces pulled from production makes it production monitoring.Why is that wrong?

      evaluate() can score existing traces, including production ones, but it is still an on-demand batch evaluation run. Monitoring is the continuous, background scoring of sampled live traces by scheduled scorers.

      Covered in Running the evaluation phase with mlflow.genai.evaluate()

    2. 2.A scorer written against an answer sheet always works unchanged in production monitoring.Why is that wrong?

      If the answer-sheet traces differ from what production emits, the scorer's parsing may not match live traces and may need rewriting.

      Covered in Running the evaluation phase with mlflow.genai.evaluate()

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “evaluation happens during development and monitoring happens once the application is deployed to production, but the fundamental components are similar”
      ↩︎ Two phases of one quality loop
      “it is not feasible to read every single response each time you evaluate to determine if the output is correct”
      ↩︎ What evaluation measures: scorers and judges
      “a curated set of evaluation queries (and ideally outputs) that are representative of the application's intended use”
      ↩︎ Running the evaluation phase with mlflow.genai.evaluate()
    2. 2.
      “Improving an agent isn't a one-time step — it's a loop you run continuously.”
      ↩︎ Two phases of one quality loop
      “Evaluate the fix. Re-run evaluation against your dataset to confirm the fix — and that you didn't break anything else.”
      ↩︎ Checkpoint
    3. 3.
      “Evaluation measures whether your agent is actually good — is the response correct, grounded, safe, and complete?”
      ↩︎ What evaluation measures: scorers and judges
      “Start by scoring a few traces from the UI, then move the same scorers into a repeatable mlflow.genai.evaluate() loop as your app matures.”
      ↩︎ Running the evaluation phase with mlflow.genai.evaluate()
    4. 4.
      “Returns the quality assessment as Feedback to attach to the trace”
      ↩︎ What evaluation measures: scorers and judges
      “You can use the same scorer for evaluation in development and monitoring in production to keep evaluation consistent throughout the application lifecycle.”
      ↩︎ Key concept
      “A scorer receives a Trace from either evaluate() or the monitoring service.”
      ↩︎ Checkpoint
    5. 5.
      “Validating prompt or model changes across app versions”
      ↩︎ Running the evaluation phase with mlflow.genai.evaluate()
      “You provide the inputs and the output, and evaluate() runs scorers and logs an evaluation run.”
      ↩︎ Exam trap 1
      “you may need to re-write your scorer functions to use them for production monitoring”
      ↩︎ Exam trap 2
      “Evaluation data can consist of existing traces, or of inputs and pre-computed outputs.”
      ↩︎ Prediction
      “Before a release or PR to prevent quality regressions”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    GenAI Production Monitoring vs Evaluation in MLflow

    Spotted a mistake, or was something unclear? Tell us.