CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 24/56

    Comparing GenAI App Versions with MLflow Evaluation Runs

    Select the best model for a given task based on common metrics generated in experiments

    13 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why an MLflow evaluation run is the evidence used to choose between candidate models, prompts or app versions
    • Set up a fair comparison by keeping the evaluation dataset and scorers fixed while changing one candidate
    • Compare evaluation runs in the MLflow UI and in code, and spot per-metric regressions that an overall average hides

    Key concept

    Evaluation run comparison — You pick the best model or app version by running each candidate through mlflow.genai.evaluate() on the same dataset with the same scorers. Then you compare the stored metrics of those runs to see where quality improved and where it regressed.

    1.Why model selection starts with an evaluation run

    Choosing the "best" model for a GenAI task on Databricks comes down to metrics that your own experiments produced, on your own task. In MLflow 3 those metrics come from an evaluation run: the result of calling mlflow.genai.evaluate() with an evaluation dataset and a list of scorers.

    An evaluation run does three things. It runs your app on every input in the dataset and captures a trace for each one. It applies each scorer to every trace, which produces feedback assessments. It then stores aggregated pass rates and metrics alongside the individual traces. You get one summary number per scorer, and behind each number are the exact traces that produced it.

    Each scorer becomes a metric you can compare. The built-in LLM judges are pre-built evaluators for common dimensions such as correctness, relevance, safety, groundedness and guideline adherence. Custom LLM judges apply your own rubric. Code-based scorers are deterministic Python checks, for example exact match, format validation, latency checks or business rules.

    Model selection then works like this: each candidate gets its own evaluation run, whether it's a different LLM, a different prompt version or a different retrieval setup. Choosing between candidates means comparing those runs. Evaluation runs are a special type of MLflow Run, so you can compare them in the UI and also query them programmatically.

    Which metrics matter depends on what the task needs. Databricks groups the common metrics for a RAG application into retrieval quality, response quality, and system performance (cost and latency). Response metrics cover whether the answer is correct, relevant, grounded in the retrieved context (not a hallucination) and safe. Cost and latency are computed deterministically from the app's outputs, as token counts and execution time. Collect both retrieval and response metrics, because an app can respond poorly despite retrieving the correct context, and the reverse. The table also shows which metrics need ground truth, meaning a known-good answer or document list in your evaluation set.

    Common metrics for comparing candidates, from the Databricks guidance on assessing RAG performance
    DimensionMetric nameMeasured byNeeds ground truth?
    Retrievalchunk_relevance/precisionLLM judgeNo
    Retrievaldocument_recallDeterministicYes
    Retrievalcontext_sufficiencyLLM judgeYes
    ResponsecorrectnessLLM judgeYes
    Responserelevance_to_queryLLM judgeNo
    ResponsegroundednessLLM judgeNo
    ResponsesafetyLLM judgeNo
    Costtotal_token_countDeterministicNo
    Latencylatency_secondsDeterministicNo

    Checkpoint 1 of 7· Check yourself

    A team wants evidence on whether swapping the LLM behind their RAG app improves quality. What does an evaluation run produce that they can compare across candidates?

    Sources12

    2.Change one thing: same dataset, same judges

    Comparing two runs only tells you something if the runs differ in the one thing you are choosing between. In the Databricks tutorial, the team improves an email-generation app and then reruns evaluation on the new version "using the same judges and dataset". The function being evaluated is passed as predict_fn, while the dataset and the scorers stay fixed.

    Be careful about what changes inside that function. The tutorial's v2 function rewrites the prompt and also calls databricks-claude-sonnet-4-5, so that comparison bundles changes together. That is fine for asking whether the new app version is better overall, but it cannot tell you whether the prompt or the LLM caused the difference. To pick between LLMs, vary only the model and hold the prompt fixed.

    The prompt-evaluation guide states the fairness rule as a best practice: evaluate every version against the same data. Suppose v2 scores higher on a dataset that is easier, or with a looser judge. That says nothing about whether v2 is the better model. Wrapping the call in mlflow.start_run(run_name=...) gives each candidate a readable name in the UI, which you will need when you pick runs to compare.

    Checkpoint 2 of 7· Fill the gap

    This snippet evaluates an improved app version so it can be compared with v1. Which argument name supplies the judges, which must stay the same as in the v1 run?

    with mlflow.start_run(run_name="v2"):
        eval_results_v2 = mlflow.genai.evaluate(
            data=eval_dataset, # same eval dataset
            predict_fn=generate_sales_email_v2, # new app version
             ? =email_judges,

    Sources34

    3.Reading a side-by-side comparison

    In the MLflow UI you compare evaluation runs from the experiment. Open the experiment, click Evaluation runs in the left sidebar, tick the runs you want to compare, and choose Compare from the Actions menu. The right pane then compares each trace across the selected runs. Click a request identifier to see that request's full traces from each run, and click See details to see each assessment. The tutorial calls this comparison view the primary way to verify that a new prompt version outperforms the previous one.

    Checkpoint 3 of 7· Put it in order

    The Databricks quality loop connects production monitoring to evaluation. Put its stages in order.

    1. 1.You write or tune a scorer
    2. 2.You evaluate to confirm the fix
    3. 3.You curate the failure cases into a dataset
    4. 4.Production monitoring identifies new failure cases
    Version comparison from the Databricks evaluate-and-improve tutorial (mean scores per judge)
    MetricV1 scoreV2 scoreChange
    safety1.0001.000+0.000
    professional_tone1.0001.000+0.000
    follows_instructions0.5710.714+0.143
    includes_next_steps0.2860.571+0.286
    mentions_contact_name1.0001.000+0.000
    retrieval_groundedness0.8570.571-0.286
    concise_communication0.2861.000+0.714
    relevance_to_query0.7141.000+0.286

    Quality scores are only one part of the choice. Latency and token counts (a proxy for cost) are separate metrics, and requirements often fix a limit on them. Databricks notes that most production applications have a latency budget, and that low-latency, real-time use cases differ from high-throughput batch ones. A practical way to combine the two kinds of evidence is to treat the latency or token budget as a requirement first. Then compare the quality scores of the candidates that meet it, and check those candidates for per-metric regressions. A candidate that is slightly better on quality but outside the budget does not satisfy the task.

    Checkpoint 4 of 7· Exam question

    A generative AI engineer runs an MLflow evaluation comparing two candidate models for a support-ticket summarization app. Both models pass the safety judge, and both comfortably beat the end-to-end latency budget. Model A scores 0.92 on the correctness scorer while Model B scores 0.81 but responds about twice as fast. Based on the logged experiment metrics, which model should be selected?

    Sources351

    4.Comparing evaluation metrics in code

    Evaluation runs are MLflow runs, so mlflow.search_runs() returns their metrics as DataFrame columns. The tutorial fetches each run by its run_id in a separate call, because search_runs doesn't support IN or OR operators. It then keeps only the quality metrics: columns that start with metrics. and end with /mean.

    Fetching two evaluation runs and selecting the per-judge mean metrics for comparisonpython
    import pandas as pd
    
    # Fetch runs separately since mlflow.search_runs doesn't support IN or OR operators
    run_v1_df = mlflow.search_runs(
        filter_string=f"run_id = '{eval_results_v1.run_id}'"
    )
    run_v2_df = mlflow.search_runs(
        filter_string=f"run_id = '{eval_results_v2.run_id}'"
    )
    
    # Extract metric columns (they end with /mean, not .aggregate_score)
    # Skip the agent metrics (latency, token counts) for quality comparison
    metric_cols = [col for col in run_v1_df.columns
                   if col.startswith('metrics.') and col.endswith('/mean')
                   and 'agent/' not in col]

    Notice the 'agent/' not in col filter. The evaluation run also records agent metrics such as latency and token counts. They are real experiment metrics, but they measure a different thing from the judge scores. The tutorial leaves them out of the quality average so that, for example, a faster version doesn't look "higher quality" because of its speed. When cost or speed matters for the choice, read those agent metrics as a separate criterion, for example against a latency or token budget, rather than folding them into one quality number.

    Checkpoint 5 of 7· Check yourself

    In the tutorial's comparison code, why are columns containing agent/ excluded from metric_cols?

    Not every scorer can run on any dataset. An evaluation dataset has inputs, and optionally expectations: the correct output, used by scorers that need ground truth. Judges such as Correctness compare the agent's response against a known-good answer, so they need expectations. Reference-free judges such as safety, groundedness and relevance to the query do not need a known-good answer, so they can run on inputs alone. Before you plan a model comparison, check that your dataset carries the expectations your chosen metrics need.

    Checkpoint 6 of 7· Check yourself

    You are comparing two LLMs on a dataset that has inputs but no expected answers. Which metric can you still compute?

    Checkpoint 7 of 7· Exam question

    A team is comparing three candidate models for a RAG question-answering app. They have curated 200 questions with expected answers and want to score both factual correctness and the presence of unsafe content in each model's responses. Which statement correctly matches each metric type to whether it needs the curated expected answers as ground truth?

    Sources312

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Two candidates' scores can be compared directly even if each was evaluated on its own dataset or with its own judges.Why is that wrong?

      A comparison is only fair when every version is evaluated against the same data with the same scorers. Only the candidate itself should change.

      Covered in Change one thing: same dataset, same judges

    2. 2.If the new version's overall average score is higher, it is better on every dimension and can be chosen without further checks.Why is that wrong?

      An average can hide a regression. In the tutorial, retrieval_groundedness dropped even though the overall average rose 20%. The comparison exists to catch regressions as well as improvements.

      Covered in Reading a side-by-side comparison

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Applies each scorer to every trace, producing feedback assessments.”
      ↩︎ Why model selection starts with an evaluation run
      “Pre-built LLM-powered evaluators for common dimensions: correctness, relevance, safety, groundedness, guideline adherence, and more.”
      ↩︎ Why model selection starts with an evaluation run
      “Evaluation runs are a special type of MLflow Run and are queryable programmatically.”
      ↩︎ Why model selection starts with an evaluation run
      “you curate them into a dataset → you write or tune a scorer → you evaluate to confirm the fix”
      ↩︎ Reading a side-by-side comparison
      “judges like Correctness compare the agent's response against a known-good answer”
      ↩︎ Comparing evaluation metrics in code
      “expectations (optional): the correct output, used by scorers that need ground truth”
      ↩︎ Comparing evaluation metrics in code
      “Use evaluation runs to answer: did this change improve quality? and did it regress anything else?”
      ↩︎ Key concept
      “Stores aggregated pass rates and metrics alongside the individual traces.”
      ↩︎ Checkpoint
    2. 2.
      “Overall latency and token consumption are examples of chain performance metrics.”
      ↩︎ Why model selection starts with an evaluation run
      “A RAG application can respond poorly despite retrieving the correct context; it can also provide good responses based on faulty retrievals.”
      ↩︎ Why model selection starts with an evaluation run
      “Some LLM judges, such as answer correctness, compare the human-labeled ground truth vs. the app outputs.”
      ↩︎ Comparing evaluation metrics in code
      “Other LLM judges, such as groundedness, do not require human-labeled ground truth to assess their app outputs.”
      ↩︎ Checkpoint
    3. 3.
      “Run the evaluation on the improved version using the same judges and dataset to see if you've successfully addressed the issues.”
      ↩︎ Change one thing: same dataset, same judges
      “model="databricks-claude-sonnet-4-5",”
      ↩︎ Change one thing: same dataset, same judges
      “This comparison view is the primary way to verify that a new prompt version outperforms the previous one.”
      ↩︎ Reading a side-by-side comparison
      “Look for specific examples where the evaluation metrics regressed so you can focus on those.”
      ↩︎ Reading a side-by-side comparison
      “Fetch runs separately since mlflow.search_runs doesn't support IN or OR operators”
      ↩︎ Comparing evaluation metrics in code
      “Skip the agent metrics (latency, token counts) for quality comparison”
      ↩︎ Comparing evaluation metrics in code
      “Compare versions to verify improvements worked and did not cause regressions.”
      ↩︎ Exam trap 2
    4. 4.
      “Use consistent datasets: Evaluate all versions against the same data for fair comparison.”
      ↩︎ Change one thing: same dataset, same judges
      “Use consistent datasets: Evaluate all versions against the same data for fair comparison.”
      ↩︎ Exam trap 1

    Continue to page 2 of 2

    Ranking Models by Experiment Metrics in MLflow

    Spotted a mistake, or was something unclear? Tell us.