What you will be able to do
- Filter MLflow runs by metric values and know which logged value a metric filter uses
- Use the chart view and parallel coordinates plots to see how parameters relate to metrics
- Link evaluation metrics to Logged Models and rank them with search_logged_models
- Move the top-ranked model to Unity Catalog registration and deployment review
1.Filtering experiment runs on metric values
An MLflow experiment is a collection of related runs. Each run records the parameters it used and the metrics it produced, saved as key-value pairs. Within an experiment you can compare and filter runs to see how a model performs and how that performance depends on parameter settings and input data. When there are only a few candidates, comparing runs by eye works. With dozens of runs, you narrow the field with a metric query in the search field on the experiment details page.
| Search expression | What it matches |
|---|---|
| metrics.r2 > 0.3 | Runs whose last logged r2 is above 0.3 |
| params.elasticNetParam = 0.5 AND metrics.avg_areaUnderROC > 0.3 | Runs with that parameter setting AND a last logged avg_areaUnderROC above 0.3 |
| MIN(metrics.rmse) <= 1 | Runs whose minimum logged rmse is at most 1 |
| MAX(metrics.memUsage) > 0.9 | Runs whose maximum logged memUsage exceeds 0.9 |
| tags.estimator_name="RandomForestRegressor" | Runs tagged with that estimator name (string values in quotes) |
MIN and MAX only work for runs logged after August 2024, because only those runs store minimum and maximum metric values. You can also filter by run state (Active or Deleted), creation time and the datasets used. The datasets filter is useful when you want to compare only candidates that were scored on the same data.
Checkpoint 1 of 5· Check yourself
You want to find runs whose rmse dropped to 1 or lower at any point during training, even if a later epoch logged a higher value. Which query does that?
A plain filter and LATEST() test the last logged value. MIN() tests the minimum logged value, so it matches runs whose rmse reached 1 or lower at any step.
“Using MIN or MAX lets you search for runs based on the minimum or maximum metric values, respectively.”Source: docs.databricks.com
2.Charting runs to see metric trade-offs
A filtered list tells you which runs pass a threshold. It doesn't tell you *why* some runs score better. The chart view on the experiment details page compares runs graphically. By default it shows only the most recent 10 runs, so before you conclude anything about the experiment as a whole, use the control at the top of the run list to change how many runs are displayed. You can sort runs by a parameter, group them by one or more parameter values, filter them with the same search syntax as the runs table, and add charts. In MLflow 3 the same charting is also available for models on the Models tab.
For model selection, the most useful chart is the parallel coordinates plot. You choose the parameters and metrics to investigate, and each run is drawn as a line across them. In the docs' example, the highlighted runs suggest that lower max_depth values go with higher auc. A pattern like that tells you which settings produce the metric you're optimising, as well as which single run happened to win.
Checkpoint 2 of 5· Check yourself
You open the chart view of an experiment containing 60 runs, and the best run on the chart has an accuracy of 0.91. What should you check before reporting 0.91 as the best result in the experiment?
By default the chart view plots only the 10 most recent runs, so a better older run may be missing from the chart until you change the number of runs displayed.
“By default, charts on this page show the most recent 10 runs.”Source: docs.databricks.com
Sources3
3.Ranking Logged Models by a metric
Runs are jobs. What you actually deploy is a model. MLflow 3 makes the model its own object, the LoggedModel, created by log_model() and identified by a unique model_id. For AI applications, a LoggedModel can represent a git commit or a set of parameters, and it can be linked to traces and metrics. Training runs output models. Evaluation runs take an existing model as input and produce metrics about it. To make those metrics rankable per model, pass the model's model_id when you log them:
# Log evaluation metrics and associate with agent
mlflow.log_metrics(
metrics=result.metrics,
dataset=eval_dataset,
# Specify the ID of the agent logged above
model_id=logged_model.model_id
)Once metrics are attached to models, mlflow.search_logged_models() can filter and order them. Numeric metrics accept =, !=, >, <, >= and <=. Metric filters can be scoped to a dataset by name and digest, and only models with matching metric values on those datasets are returned. That scoping matters: an accuracy measured on the training set is not the same evidence as one measured on held-out evaluation data. You can also go the other way, from a model to the runs that used it, with mlflow.search_runs(filter_string = "models.model_id = <my-model-id>").
In the Databricks deep-learning example, a checkpoint is logged every 10 epochs as its own LoggedModel. The checkpoints are then ranked by accuracy. The best is 0.955 at step 90 and the worst is 0.357 at step 0. Note that the printed metrics in that example carry dataset_name='train', so those accuracies were measured on the training dataset. Treat the example as a demonstration of the ranking API; to choose a model to deploy, scope the ranking to held-out evaluation data.
Checkpoint 3 of 5· Fill the gap
This code ranks checkpoint models so that the highest-accuracy model is at index 0. Which value completes the sort?
ranked_checkpoints = mlflow.search_logged_models(
output_format="list",
order_by=[{"field_name": "metrics.accuracy", "ascending": ? }]
)
best_checkpoint: mlflow.entities.LoggedModel = ranked_checkpoints[0]
print(best_checkpoint.metrics[0])Sorting with ascending False puts the largest accuracy first, so ranked_checkpoints[0] is the best checkpoint and ranked_checkpoints[-1] is the worst. For any metric, choose the sort direction that puts the preferred value first.
Source: docs.databricks.com4.Choosing a primary metric and promoting the winner
Every ranking above assumes you have already decided which metric decides the choice. Databricks AutoML makes that decision an explicit parameter. Its primary_metric is the metric used to evaluate and rank model performance, and each problem type has a default. The same discipline applies when you rank runs or Logged Models yourself: name the deciding metric before you look at the results.
| Problem type | Supported primary_metric values |
|---|---|
| Classification | f1 (default), log_loss, precision, accuracy, roc_auc |
| Regression | r2 (default), mae, rmse, mse |
After the winner is identified, Databricks provides two views with different jobs. The experiment's Models tab shows the logged models from one experiment on a single page. Its Charts tab helps you compare them and pick the version to register to Unity Catalog. From the model's details page you click Register model and choose Unity Catalog. Click it only once, because each click registers a duplicate. In Catalog Explorer, the model version page then gathers parameters, metrics and traces from every linked workspace, endpoint and experiment. An evaluation task in a deployment job adds further metrics there, and the job's approver reviews that page before approving the version for deployment.
Checkpoint 4 of 5· Check yourself
A deployment job has an evaluation task and requires approval. Where does the approver look to decide whether the registered model version should be deployed?
The Catalog Explorer model version page brings together metrics and traces from every linked environment, including the metrics from the deployment job's evaluation task, and the approver reviews it there.
“The approver for the job can then review this page to assess whether to approve the model version for deployment.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A team is building a retrieval system over legal contracts where individual clauses are chunked at up to 6,000 tokens each. While comparing embedding model cards to select a model for this pipeline, which approach correctly uses the context length metric to guide the decision?
Correct answer: C — Select the embedding model whose documented maximum context length exceeds the largest chunk size, since embeddings computed on truncated text would drop clause content and reduce retrieval quality
- A. Parameter count and maximum input context length are separate model card attributes; a larger embedding model is not guaranteed to accept longer sequences than its published context length allows.
- B. Inference latency doesn't address whether clause text fits within the embedding model's input window, and vector search does not automatically re-chunk oversized text to fit an embedding model's context length.
- C. Correct. If a chunk exceeds the embedding model's maximum context length, the excess text is truncated before embedding, silently dropping clause content and degrading retrieval quality, so the model's context length must cover the largest chunk size.
- D. Context length is a real attribute of embedding models, not just generation models, since embedding models also have a fixed maximum input window; a general retrieval benchmark score doesn't guarantee the model can encode a 6,000-token chunk without truncation.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A run search such as metrics.rmse <= 1 matches any run whose metric reached that value at some point during training.Why is that wrong?
Plain metric filters test the last logged value. To filter on the minimum or maximum value logged, wrap the metric in MIN() or MAX().
Covered in Filtering experiment runs on metric values
2.The experiment chart view plots every run, so the best line on the chart is the best run in the experiment.Why is that wrong?
By default the chart view shows only the 10 most recent runs. You have to change the number of runs displayed before you compare across the whole experiment.
Covered in Charting runs to see metric trade-offs
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/mlflow/trackingOfficial docs
“you can compare and filter runs to understand how your model performs”
↩︎ Filtering experiment runs on metric values - 2.https://docs.databricks.com/aws/en/mlflow/runsOfficial docs
“Only runs logged after August 2024 have minimum and maximum metric values.”
↩︎ Filtering experiment runs on metric values“By default, metric values are filtered based on the last logged value.”
↩︎ Exam trap 1“By default, metric values are filtered based on the last logged value.”
↩︎ Prediction“Using MIN or MAX lets you search for runs based on the minimum or maximum metric values, respectively.”
↩︎ Checkpoint - 3.
“A parallel coordinates plot is useful in understanding the effect of parameter settings on model performance and investigating relationships between parameters and metrics.”
↩︎ Charting runs to see metric trade-offs“the runs highlighted in the black boxes suggest that lower values for max_depth result in higher values for the metric auc.”
↩︎ Charting runs to see metric trade-offs“By default, charts on this page show the most recent 10 runs.”
↩︎ Exam trap 2“By default, charts on this page show the most recent 10 runs.”
↩︎ Checkpoint - 4.
“Training runs produce models as outputs, and evaluation runs use existing models as input to produce metrics”
↩︎ Ranking Logged Models by a metric“Logged Model tracking lets you compare models against each other, find the most performant model, and track down information during debugging.”
↩︎ Ranking Logged Models by a metric“You can filter metrics based on dataset-specific performance, and only models with matching metric values on the given datasets are returned.”
↩︎ Ranking Logged Models by a metric - 5.
“dataset_name='train',”
↩︎ Ranking Logged Models by a metric“provides visualizations to help you compare models and select the model versions to register to Unity Catalog”
↩︎ Choosing a primary metric and promoting the winner“The approver for the job can then review this page to assess whether to approve the model version for deployment.”
↩︎ Checkpoint - 6.
“Metric used to evaluate and rank model performance.”
↩︎ Choosing a primary metric and promoting the winner