What you will be able to do
- Explain what an MLflow run records and why runs are the unit you rank to find the best model
- Write a search_runs call that orders runs by a metric, breaks ties, and takes the top result
- Narrow the candidates with filter expressions on metrics, params and tags, including MIN/MAX and their limits
- Turn the best run's run_id into a runs:/ model URI you can load or register
- Rank MLflow 3 Logged Models with search_logged_models and tell it apart from search_runs
Key concept
Run as the unit of comparison — Each training execution is logged as a run with its own parameters, metrics and tags. To find the best model, you search the runs in an experiment, sort them by the metric you care about, and take the run_id of the top result.
1.What you are ranking: runs inside an experiment
"Best run" only makes sense once you know what a run holds. MLflow tracking organises model development into experiments and runs. A run is one execution of your training code. An experiment groups related runs, and that grouping is what lets you compare them side by side.
Each run stores the values you need to rank it: model parameters and metrics as key-value pairs, tags as free-form metadata, and any artifacts the run produced. Finding the best run means asking a question of these stored values, for example: which run in this experiment has the highest test AUC? Runs are the rows, and their metrics and params are the columns you sort and filter on.
You also need to know where your runs were logged. Every run goes to the active experiment. If you never set one, Databricks logs runs to the notebook experiment, which is the experiment attached to the notebook that ran the code. A search that points at a different experiment will not find them.
Checkpoint 1 of 8· Check yourself
A data scientist trains several models in a Databricks notebook and never calls mlflow.set_experiment. Where are the runs logged?
If no experiment has been set as active, runs fall back to the notebook experiment. A search for the best run has to look in that experiment.
“If you have not explicitly set an experiment as the active experiment, runs are logged to the notebook experiment.”Source: docs.databricks.com
2.Ranking runs with search_runs and order_by
The MLflow API indexes mlflow.search_runs() as the call that searches and filters runs by criteria. The Databricks getting-started tutorial uses it to pick the best of a hyperparameter sweep. The sample below is the pattern the exam expects you to recognise.
# Sort runs by their test auc. In case of ties, use the most recent run.
best_run = mlflow.search_runs(
order_by=['metrics.test_auc DESC', 'start_time DESC'],
max_results=10,
).iloc[0]
print('Best Run')
print('AUC: {}'.format(best_run["metrics.test_auc"]))
print('Num Estimators: {}'.format(best_run["params.n_estimators"]))
print('Max Depth: {}'.format(best_run["params.max_depth"]))
print('Learning Rate: {}'.format(best_run["params.learning_rate"]))Here is what each part does:
- **order_by** takes a list of sort keys. Metrics are referenced with the metrics. prefix, and DESC puts the highest AUC first. For a metric where lower is better, you would reverse the direction.
- **The second key, start_time DESC, decides between runs with the same AUC. The comment says why: in case of ties, use the most recent run.
- max_results=10 limits how many runs the search returns. You only need the top one.
- .iloc[0] takes the first row of the sorted result. That row is the best run.
- Column names mirror the prefixes:** the winner's values are read as best_run["metrics.test_auc"] and best_run["params.max_depth"].
A note on scope: the exam objective says "MLflow Client API". The supplied Databricks documentation shows this ranking only through mlflow.search_runs. It does not include an example that creates an MlflowClient object, so this lesson does not describe that class's method signature.
Checkpoint 2 of 8· Check yourself
In the tutorial's search_runs call, what does the second sort key 'start_time DESC' do?
Sort keys are applied in order. test_auc decides the ranking, and start_time only matters between runs whose AUC is equal.
“In case of ties, use the most recent run.”Source: docs.databricks.com
Checkpoint 3 of 8· Exam question
A data scientist logged 40 runs of a regression model to an MLflow experiment on Databricks, recording `rmse` as a metric on every run. They want the run with the lowest `rmse` value returned directly as a single result, without pulling all 40 runs into a DataFrame and sorting them in pandas. Which approach identifies that run using the MLflow Client API?
Correct answer: A — Call `MlflowClient().search_runs(experiment_ids=[exp_id], order_by=["metrics.rmse ASC"], max_results=1)` and read the first `Run` object returned in the list.
- A. Sorting with `order_by=["metrics.rmse ASC"]` and `max_results=1` asks the tracking server to rank runs from lowest to highest rmse and return only the top match, which is exactly the lowest-error run without any client-side sorting.
- B. `DESC` sorts from highest to lowest rmse, so `max_results=1` would return the run with the worst error instead of the best one.
- C. `list_experiments()` returns experiment-level metadata such as names and artifact locations; metrics like `rmse` are logged per run, not stored as tags on the experiment itself.
- D. This requires already knowing every run ID in advance and issuing one call per run, which is precisely the manual full-scan approach that `order_by` and `max_results` are designed to replace.
3.Narrowing the candidates before you rank
Sorting the whole experiment is not always enough. You may only want runs from one estimator, or runs that pass a quality bar. You narrow the set with filter expressions on metrics, params and tags. The runs documentation describes this syntax for the experiment page's search box. The Logged Models documentation points to the same "filtering for runs" syntax for the API's filter_string, which it uses in mlflow.search_runs(filter_string = ...).
| Expression | What it matches |
|---|---|
| metrics.r2 > 0.3 | Runs whose r2 metric (last logged value) exceeds 0.3 |
| params.elasticNetParam = 0.5 AND metrics.avg_areaUnderROC > 0.3 | Runs with that parameter value and an AUC above 0.3 |
| MIN(metrics.rmse) <= 1 | Runs whose minimum logged rmse is at most 1 |
| LATEST(metrics.memUsage) = 0 AND MIN(metrics.rmse) <= 1 | Combines the last value of one metric with the minimum of another |
| tags.estimator_name="RandomForestRegressor" | Runs tagged with that estimator; string values need quotes |
tags.my custom tag = "my value" | A tag key that contains spaces, wrapped in backticks |
This default catches people out when a metric is logged at every step. A plain metrics.x filter compares the final value. To filter on the best or worst value seen, wrap the metric in MIN(...) or MAX(...). These aggregates have a time limit, though: only runs logged after August 2024 store minimum and maximum metric values. Older runs will not match a MIN or MAX filter, even if their history would have qualified.
Checkpoint 4 of 8· Check yourself
A filter MIN(metrics.rmse) <= 1 returns recent runs but misses a 2023 run that you know reached rmse 0.8. What is the most likely reason?
MIN and MAX rely on stored aggregate values, and runs logged before August 2024 do not have them.
“Only runs logged after August 2024 have minimum and maximum metric values.”Source: docs.databricks.com
Checkpoint 5 of 8· Exam question
A team logged runs under three model families using a `model_type` tag with values `xgboost`, `random_forest`, and `logistic_regression`, all in the same MLflow experiment. A data scientist needs the single best `xgboost` run, ranked by `accuracy`, while ignoring the other two families entirely. Which call to `search_runs` returns exactly that run?
Correct answer: A — `search_runs(experiment_ids=[exp_id], filter_string="tags.model_type = 'xgboost'", order_by=["metrics.accuracy DESC"], max_results=1)`
- A. The filter string scopes the search to runs tagged `xgboost` before any sorting happens, and `DESC` on accuracy then puts the highest-accuracy run in that scoped set first.
- B. `model_type` was logged as a tag, not a metric, so referencing it as `metrics.model_type` in the filter string will not match any run and the search returns an empty result.
- C. Sorting by the tag first ranks runs alphabetically by model family before accuracy is even considered, and it never actually restricts the result set to only xgboost runs.
- D. `ASC` ordering ranks the lowest-accuracy xgboost run first, which is the opposite of the best-performing run the scenario asks for.
4.From the best run's ID to a usable model
The search returns the winning run's metrics and params. What you usually need next is its model, and the link between the two is run_id. The tutorial builds a runs:/<run_id>/model URI from the best run and loads the model through it.
best_model_pyfunc = mlflow.pyfunc.load_model(
'runs:/{run_id}/model'.format(
run_id=best_run.run_id
)
)The tutorial uses the same URI as the input to mlflow.register_model, with a three-level Unity Catalog name. Registration belongs to a separate objective. The point here is that the best-run search ends in a run_id, and every later step depends on that ID.
Checkpoint 6 of 8· Fill the gap
Which scheme completes the URI that points at the best run's logged model?
model_uri = ' ? :/{run_id}/model'.format(
run_id=best_run.run_id
)A runs:/ URI identifies a model by the run that logged it. That is why the run_id returned by search_runs is all you need.
Source: docs.databricks.comSources4
5.MLflow 3: ranking Logged Models instead of runs
MLflow 3 makes a model its own tracked entity, separate from the run that produced it. It is no longer just a run artifact, and you can rank models directly. The deep learning workflow page ranks checkpoint models by accuracy with mlflow.search_logged_models.
ranked_checkpoints = mlflow.search_logged_models(
output_format="list",
order_by=[{"field_name": "metrics.accuracy", "ascending": False}]
)
best_checkpoint: mlflow.entities.LoggedModel = ranked_checkpoints[0]
print(best_checkpoint.metrics[0])The two APIs connect. Each metric on a Logged Model also records the run_id that produced it. Going the other way, mlflow.search_runs(filter_string = "models.model_id = <my-model-id>") returns all runs that used a given model as an input or output. The table below shows the syntax differences that exam distractors tend to swap.
| Aspect | mlflow.search_runs | mlflow.search_logged_models |
|---|---|---|
| What it ranks | Runs (single executions of training code) | Logged Models (MLflow 3 model entities) |
| order_by form | Strings such as 'metrics.test_auc DESC' | Dicts such as {"field_name": "metrics.accuracy", "ascending": False} |
| Taking the top result | best_run = ... .iloc[0] | ranked_checkpoints[0] with output_format="list" |
| Identifier for the next step | best_run.run_id | model_id (each Metric also carries run_id) |
Checkpoint 7 of 8· Match them up
Match each expression to what it does
Tap a term, then the definition that fits it.
search_runs sorts with strings and search_logged_models sorts with dicts. A models.model_id filter goes from a model back to its runs, and a runs:/ URI goes from a run to its model.
“You can search runs by model ID to return all runs that have the Logged Model as an input or output.”Source: docs.databricks.com
Checkpoint 8 of 8· Exam question
A notebook already imports `pandas`, and the data scientist wants to quickly sort completed runs by validation F1 score inside a DataFrame using `.sort_values()` and `.head()`, rather than working with a list of `Run` objects. Which approach fits that exploratory workflow?
Correct answer: A — Call the module-level `mlflow.search_runs(experiment_ids=[exp_id])` function, which returns a pandas DataFrame with rows per run and columns like `metrics.f1`.
- A. The module-level `mlflow.search_runs` function is documented to return a pandas DataFrame with one row per run, which is exactly the shape needed for `.sort_values()` and `.head()` style exploration.
- B. The `MlflowClient` class's `search_runs` method returns a list of `Run` objects rather than a DataFrame, so pandas methods like `.sort_values()` cannot be called on it without extra conversion.
- C. `get_experiment` returns a single `Experiment` metadata object describing the experiment itself, such as its name and artifact location, not a table of per-run metrics.
- D. `list_run_infos` returns a list of `RunInfo` objects containing only run metadata like status and timestamps, not a DataFrame and not logged metric values.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Once the runs are sorted by test_auc, .iloc[0] is the best run no matter which direction you sorted.Why is that wrong?
The first row is the best run only if the sort direction matches the metric. The tutorial sorts AUC with DESC so the highest value comes first. Sorting the other way would put the worst run at index 0.
Covered in Ranking runs with search_runs and order_by
2.A filter like metrics.rmse <= 1 matches any run whose rmse was at or below 1 at some point during training.Why is that wrong?
A plain metric filter compares only the last logged value. To filter on the best value, you need MIN or MAX, and those work only for runs logged after August 2024.
Covered in Narrowing the candidates before you rank
3.If no experiment is set, runs are not tracked, so search_runs has nothing to rank.Why is that wrong?
Every run is logged to the active experiment. If none is set, runs go to the notebook experiment.
4.In MLflow 3 a model is still just a file in the run's artifacts, so search_runs is the only way to rank models.Why is that wrong?
MLflow 3 makes the model a first-class entity, and search_logged_models can rank models directly by their metrics.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/mlflow/trackingOfficial docs
“In an experiment, you can compare and filter runs to understand how your model performs”
↩︎ What you are ranking: runs inside an experiment“With MLflow 3, LoggedModels elevates the concept of a model produced by a run, establishing it as a distinct entity”
↩︎ MLflow 3: ranking Logged Models instead of runs“A run is a single execution of model code.”
↩︎ Key concept - 2.https://docs.databricks.com/aws/en/mlflow/runsOfficial docs
“model parameters and metrics saved as key-value pairs, tags for run metadata, and any artifacts, or output files, created by the run.”
↩︎ What you are ranking: runs inside an experiment“By default, metric values are filtered based on the last logged value.”
↩︎ Narrowing the candidates before you rank“String values must be enclosed in quotes as shown.”
↩︎ Narrowing the candidates before you rank“Using MIN or MAX lets you search for runs based on the minimum or maximum metric values, respectively.”
↩︎ Exam trap 2“All MLflow runs are logged to the active experiment.”
↩︎ Exam trap 3“In MLflow 3, models are now their distinct first-class object rather than being logged as run artifact.”
↩︎ Exam trap 4“If you have not explicitly set an experiment as the active experiment, runs are logged to the notebook experiment.”
↩︎ Checkpoint“Only runs logged after August 2024 have minimum and maximum metric values.”
↩︎ Checkpoint - 3.
“mlflow.search_runs() - Search and filter runs by criteria”
↩︎ Ranking runs with search_runs and order_by - 4.
“In case of ties, use the most recent run.”
↩︎ Ranking runs with search_runs and order_by“'runs:/{run_id}/model'.format(”
↩︎ From the best run's ID to a usable model“order_by=['metrics.test_auc DESC', 'start_time DESC'],”
↩︎ Exam trap 1 - 5.
“For more information on the filter string syntax, see filtering for runs.”
↩︎ Narrowing the candidates before you rank“You can search runs by model ID to return all runs that have the Logged Model as an input or output.”
↩︎ Checkpoint - 6.
“The following code shows how to rank the checkpoint models by accuracy.”
↩︎ MLflow 3: ranking Logged Models instead of runs