CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 3 · Lesson 24/56

    Ranking Models by Experiment Metrics in MLflow

    Select the best model for a given task based on common metrics generated in experiments

    9 min read
    1.79% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Filter MLflow runs by metric values and know which logged value a metric filter uses
    • Use the chart view and parallel coordinates plots to see how parameters relate to metrics
    • Link evaluation metrics to Logged Models and rank them with search_logged_models
    • Move the top-ranked model to Unity Catalog registration and deployment review

    1.Filtering experiment runs on metric values

    An MLflow experiment is a collection of related runs. Each run records the parameters it used and the metrics it produced, saved as key-value pairs. Within an experiment you can compare and filter runs to see how a model performs and how that performance depends on parameter settings and input data. When there are only a few candidates, comparing runs by eye works. With dozens of runs, you narrow the field with a metric query in the search field on the experiment details page.

    Run search expressions from the Databricks docs and what each one matches
    Search expressionWhat it matches
    metrics.r2 > 0.3Runs whose last logged r2 is above 0.3
    params.elasticNetParam = 0.5 AND metrics.avg_areaUnderROC > 0.3Runs with that parameter setting AND a last logged avg_areaUnderROC above 0.3
    MIN(metrics.rmse) <= 1Runs whose minimum logged rmse is at most 1
    MAX(metrics.memUsage) > 0.9Runs whose maximum logged memUsage exceeds 0.9
    tags.estimator_name="RandomForestRegressor"Runs tagged with that estimator name (string values in quotes)

    MIN and MAX only work for runs logged after August 2024, because only those runs store minimum and maximum metric values. You can also filter by run state (Active or Deleted), creation time and the datasets used. The datasets filter is useful when you want to compare only candidates that were scored on the same data.

    Checkpoint 1 of 5· Check yourself

    You want to find runs whose rmse dropped to 1 or lower at any point during training, even if a later epoch logged a higher value. Which query does that?

    Sources12

    2.Charting runs to see metric trade-offs

    A filtered list tells you which runs pass a threshold. It doesn't tell you *why* some runs score better. The chart view on the experiment details page compares runs graphically. By default it shows only the most recent 10 runs, so before you conclude anything about the experiment as a whole, use the control at the top of the run list to change how many runs are displayed. You can sort runs by a parameter, group them by one or more parameter values, filter them with the same search syntax as the runs table, and add charts. In MLflow 3 the same charting is also available for models on the Models tab.

    For model selection, the most useful chart is the parallel coordinates plot. You choose the parameters and metrics to investigate, and each run is drawn as a line across them. In the docs' example, the highlighted runs suggest that lower max_depth values go with higher auc. A pattern like that tells you which settings produce the metric you're optimising, as well as which single run happened to win.

    Checkpoint 2 of 5· Check yourself

    You open the chart view of an experiment containing 60 runs, and the best run on the chart has an accuracy of 0.91. What should you check before reporting 0.91 as the best result in the experiment?

    Sources3

    3.Ranking Logged Models by a metric

    Runs are jobs. What you actually deploy is a model. MLflow 3 makes the model its own object, the LoggedModel, created by log_model() and identified by a unique model_id. For AI applications, a LoggedModel can represent a git commit or a set of parameters, and it can be linked to traces and metrics. Training runs output models. Evaluation runs take an existing model as input and produce metrics about it. To make those metrics rankable per model, pass the model's model_id when you log them:

    Linking evaluation metrics to a LoggedModel via model_idpython
      # Log evaluation metrics and associate with agent
      mlflow.log_metrics(
        metrics=result.metrics,
        dataset=eval_dataset,
        # Specify the ID of the agent logged above
        model_id=logged_model.model_id
      )

    Once metrics are attached to models, mlflow.search_logged_models() can filter and order them. Numeric metrics accept =, !=, >, <, >= and <=. Metric filters can be scoped to a dataset by name and digest, and only models with matching metric values on those datasets are returned. That scoping matters: an accuracy measured on the training set is not the same evidence as one measured on held-out evaluation data. You can also go the other way, from a model to the runs that used it, with mlflow.search_runs(filter_string = "models.model_id = <my-model-id>").

    In the Databricks deep-learning example, a checkpoint is logged every 10 epochs as its own LoggedModel. The checkpoints are then ranked by accuracy. The best is 0.955 at step 90 and the worst is 0.357 at step 0. Note that the printed metrics in that example carry dataset_name='train', so those accuracies were measured on the training dataset. Treat the example as a demonstration of the ranking API; to choose a model to deploy, scope the ranking to held-out evaluation data.

    Checkpoint 3 of 5· Fill the gap

    This code ranks checkpoint models so that the highest-accuracy model is at index 0. Which value completes the sort?

    ranked_checkpoints = mlflow.search_logged_models(
      output_format="list",
      order_by=[{"field_name": "metrics.accuracy", "ascending":  ? }]
    )
    
    best_checkpoint: mlflow.entities.LoggedModel = ranked_checkpoints[0]
    print(best_checkpoint.metrics[0])

    Sources45

    4.Choosing a primary metric and promoting the winner

    Every ranking above assumes you have already decided which metric decides the choice. Databricks AutoML makes that decision an explicit parameter. Its primary_metric is the metric used to evaluate and rank model performance, and each problem type has a default. The same discipline applies when you rank runs or Logged Models yourself: name the deciding metric before you look at the results.

    Supported primary_metric values in the AutoML API (defaults marked)
    Problem typeSupported primary_metric values
    Classificationf1 (default), log_loss, precision, accuracy, roc_auc
    Regressionr2 (default), mae, rmse, mse

    After the winner is identified, Databricks provides two views with different jobs. The experiment's Models tab shows the logged models from one experiment on a single page. Its Charts tab helps you compare them and pick the version to register to Unity Catalog. From the model's details page you click Register model and choose Unity Catalog. Click it only once, because each click registers a duplicate. In Catalog Explorer, the model version page then gathers parameters, metrics and traces from every linked workspace, endpoint and experiment. An evaluation task in a deployment job adds further metrics there, and the job's approver reviews that page before approving the version for deployment.

    Checkpoint 4 of 5· Check yourself

    A deployment job has an evaluation task and requires approval. Where does the approver look to decide whether the registered model version should be deployed?

    Checkpoint 5 of 5· Exam question

    A team is building a retrieval system over legal contracts where individual clauses are chunked at up to 6,000 tokens each. While comparing embedding model cards to select a model for this pipeline, which approach correctly uses the context length metric to guide the decision?

    Sources65

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A run search such as metrics.rmse <= 1 matches any run whose metric reached that value at some point during training.Why is that wrong?

      Plain metric filters test the last logged value. To filter on the minimum or maximum value logged, wrap the metric in MIN() or MAX().

      Covered in Filtering experiment runs on metric values

    2. 2.The experiment chart view plots every run, so the best line on the chart is the best run in the experiment.Why is that wrong?

      By default the chart view shows only the 10 most recent runs. You have to change the number of runs displayed before you compare across the whole experiment.

      Covered in Charting runs to see metric trade-offs

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “you can compare and filter runs to understand how your model performs”
      ↩︎ Filtering experiment runs on metric values
    2. 2.
      “Only runs logged after August 2024 have minimum and maximum metric values.”
      ↩︎ Filtering experiment runs on metric values
      “By default, metric values are filtered based on the last logged value.”
      ↩︎ Exam trap 1
      “By default, metric values are filtered based on the last logged value.”
      ↩︎ Prediction
      “Using MIN or MAX lets you search for runs based on the minimum or maximum metric values, respectively.”
      ↩︎ Checkpoint
    3. 3.
      “A parallel coordinates plot is useful in understanding the effect of parameter settings on model performance and investigating relationships between parameters and metrics.”
      ↩︎ Charting runs to see metric trade-offs
      “the runs highlighted in the black boxes suggest that lower values for max_depth result in higher values for the metric auc.”
      ↩︎ Charting runs to see metric trade-offs
      “By default, charts on this page show the most recent 10 runs.”
      ↩︎ Exam trap 2
      “By default, charts on this page show the most recent 10 runs.”
      ↩︎ Checkpoint
    4. 4.
      “Training runs produce models as outputs, and evaluation runs use existing models as input to produce metrics”
      ↩︎ Ranking Logged Models by a metric
      “Logged Model tracking lets you compare models against each other, find the most performant model, and track down information during debugging.”
      ↩︎ Ranking Logged Models by a metric
      “You can filter metrics based on dataset-specific performance, and only models with matching metric values on the given datasets are returned.”
      ↩︎ Ranking Logged Models by a metric
    5. 5.
      “dataset_name='train',”
      ↩︎ Ranking Logged Models by a metric
      “provides visualizations to help you compare models and select the model versions to register to Unity Catalog”
      ↩︎ Choosing a primary metric and promoting the winner
      “The approver for the job can then review this page to assess whether to approve the model version for deployment.”
      ↩︎ Checkpoint

    Spotted a mistake, or was something unclear? Tell us.