What you will be able to do
- Explain what an MLflow run records and which experiment it logs to
- Use mlflow.start_run, mlflow.log_param, mlflow.log_metric and mlflow.log_artifact to record results that autologging misses
- Combine Databricks Autologging with manual logging in the same run
- Say where each kind of logged data shows up on the run page
Key concept
MLflow run — A run is one execution of your training code. Everything you log while it is active (parameters, metrics, tags, artifacts and models) is attached to that run, and the run belongs to an experiment.
1.Runs, experiments and the Tracking API
Model development is iterative. You try a hyperparameter, check a score, change something and try again. MLflow tracking keeps a record of each attempt. It uses three concepts. A run is a single execution of your model code. An experiment is a collection of related runs, which you can compare and filter. A model is a collection of artifacts that represent a trained model.
The Tracking API is how data gets into a run. It logs parameters, metrics, tags and artifacts, and it talks to a tracking server. On Databricks that server is hosted for you, so no setup is needed. MLflow comes pre-installed on Databricks Runtime ML clusters. On a standard Databricks Runtime cluster you have to install the mlflow library yourself.
Two settings decide where your logging lands. The tracking URI picks the server, and it defaults to the current workspace. The experiment picks which experiment on that server receives the run. When you call mlflow.start_run() in a notebook, metrics and parameters go to the active experiment. If no experiment is active, Databricks creates a notebook experiment that shares the notebook's name and ID. Only runs started from inside a notebook can be logged to that notebook's experiment. To send runs from any notebook or API call to one shared place, set a workspace experiment first:
import mlflow
# By default MLflow logs to the Databricks-hosted workspace tracking server. You can connect to a different server using the tracking URI.
mlflow.set_tracking_uri("databricks://remote-workspace-url")
# Set experiment in the tracking server
mlflow.set_experiment("/Shared/my-experiment")Checkpoint 1 of 7· Check yourself
A job started through the API, not from a notebook, needs to log its runs next to runs from several notebooks. What must the code do?
Runs launched from the APIs or from any notebook can go to a workspace experiment once it is set with mlflow.set_experiment(). Only runs started inside a notebook can use that notebook's experiment.
“MLflow runs launched from any notebook or from the APIs can be logged to a workspace experiment.”Source: docs.databricks.com
2.Logging metrics and artifacts by hand
Autologging is the easiest way to start. MLflow records parameters, training metrics and the model for many frameworks without extra code. It does not know what you do after training, though. For example, it has no record of a score on a held-out test set or a plot you made yourself. For anything like that, use the MLflow logging API directly. The documentation recommends it when you want control over what is logged, or to add artifacts such as CSV files or plots.
The Databricks getting-started notebook shows the pattern. Autologging is switched on with mlflow.sklearn.autolog(), and a classifier is trained inside a named run. After that, two results are logged manually. The test AUC is a single number, so it goes in with mlflow.log_metric(key, value). The ROC curve is an image, so it is saved to a local file first and then uploaded with mlflow.log_artifact(path).
with mlflow.start_run(run_name='gradient_boost') as run:
model = sklearn.ensemble.GradientBoostingClassifier(random_state=0)
# Models, parameters, and training metrics are tracked automatically
model.fit(X_train, y_train)
predicted_probs = model.predict_proba(X_test)
roc_auc = sklearn.metrics.roc_auc_score(y_test, predicted_probs[:,1])
roc_curve = sklearn.metrics.RocCurveDisplay.from_estimator(model, X_test, y_test)
# Save the ROC curve plot to a file
roc_curve.figure_.savefig("roc_curve.png")
# The AUC score on test data is not automatically logged, so log it manually
mlflow.log_metric("test_auc", roc_auc)
# Log the ROC curve image file as an artifact
mlflow.log_artifact("roc_curve.png")| Call | What it does in the example |
|---|---|
| mlflow.sklearn.autolog() | Turns on automatic logging of the model, parameters and training metrics |
| mlflow.start_run(run_name='gradient_boost') | Opens a named run. The with block ends it automatically. |
| mlflow.log_metric("test_auc", roc_auc) | Records one numeric result, the test AUC, under a key |
| mlflow.log_artifact("roc_curve.png") | Uploads a file that already exists on local disk into the run's artifacts |
Checkpoint 2 of 7· Check yourself
A model is trained with mlflow.sklearn.autolog() enabled. Afterwards you compute ROC AUC on X_test. What do you have to do for that score to show up on the run?
Autologging records training metrics, but a test-set score you compute yourself is not captured. You log it as a metric by hand.
“The AUC score on test data is not automatically logged, so log it manually”Source: docs.databricks.com
Checkpoint 3 of 7· Exam question
A data scientist is training a scikit-learn classifier over 20 epochs inside a `with mlflow.start_run():` block on Databricks and wants the MLflow UI to plot how validation accuracy changes across epochs, not just show a single final number. Which manual logging call, placed inside the training loop, achieves this?
Correct answer: A — Call `mlflow.log_metric("val_accuracy", value, step=epoch)` once per epoch, passing the epoch number as `step` so the metric history builds one point per epoch.
- A. This is correct because passing an explicit `step` value on each call is exactly how MLflow builds a per-step metric history for a key, which the UI then renders as a trend line across epochs.
- B. This is incorrect because omitting `step` logs every call at the default step of 0, so repeated calls create multiple points stacked at the same step rather than a clean per-epoch trend.
- C. This is incorrect because parameters are meant for static configuration values, are not designed to accept repeated updates to the same key across a run, and are not charted as a time series in the UI.
- D. This is incorrect because uploaded files are stored as opaque artifacts; MLflow does not parse artifact contents and turn them into interactive metric charts.
3.Adding manual logging to an autologged run
No, it won't. Databricks Autologging does not apply to runs you create with the fluent API's mlflow.start_run(). To get autologged content into a run you opened yourself, call mlflow.autolog() explicitly. To mix both kinds of logging, Databricks gives this recipe:
1. Call mlflow.autolog() with exclusive=False.
2. Start a run with mlflow.start_run(), ideally as a with block.
3. Log pre-training content such as mlflow.log_param().
4. Train the model.
5. Log post-training content such as mlflow.log_metric().
6. If you did not use with, close the run with mlflow.end_run().
Checkpoint 4 of 7· Fill the gap
Which argument lets your own log_param and log_metric calls go into the same run as the autologged content?
import mlflow
mlflow.autolog( ? =False)
with mlflow.start_run():
mlflow.log_param("example_param", "example_value")
# <your model training code here>
mlflow.log_metric("example_metric", 5)The documented first step is to call mlflow.autolog() with exclusive=False so that autologged and manually logged content share the run. disable=True would switch autologging off.
Source: docs.databricks.comCheckpoint 5 of 7· Put it in order
Put the steps for tracking extra content in an autologged run in order
- 1.Track post-training content with mlflow.log_metric()
- 2.Start an MLflow run using mlflow.start_run()
- 3.Call mlflow.autolog() with exclusive=False
- 4.Track pre-training content with mlflow.log_param()
- 5.End the run with mlflow.end_run() if you did not use with
- 6.Train the model in a framework supported by autologging
Autologging is configured before the run starts. Parameters come before training and metrics after it. An explicit end_run is needed only when start_run was not used as a context manager.
“If you did not use with mlflow.start_run() in Step 2, end the MLflow run using mlflow.end_run().”Source: docs.databricks.com
Checkpoint 6 of 7· Exam question
After fitting a `RandomForestClassifier` with `n_estimators=200`, `max_depth=8`, and `min_samples_leaf=4`, a Databricks user wants to record all three hyperparameters to the active MLflow run with a single API call instead of three separate calls. Which call does this correctly?
Correct answer: A — `mlflow.log_params({"n_estimators": 200, "max_depth": 8, "min_samples_leaf": 4})`, passing a dictionary so all three key-value pairs are written to the run in one batched call.
- A. This is correct because `log_params` is built specifically to batch multiple hyperparameter key-value pairs into the active run in one call, keeping each hyperparameter individually searchable.
- B. This is incorrect because tuning configuration like `n_estimators` and `max_depth` are hyperparameters, not measured outcomes, so logging them as metrics misrepresents what they track and how the UI groups them.
- C. This is incorrect because it collapses three distinct, individually searchable hyperparameters into a single opaque string value under one key, losing the ability to filter or compare runs by each hyperparameter.
- D. This is incorrect because `log_artifact` expects a path to a local file on disk, not a JSON string, and the resulting file would not be indexed as searchable run parameters.
Sources4
4.Where logged data shows up, and what limits apply
Each kind of call ends up in a specific place on the run. A run records the notebook that launched it, any models it created, parameters and metrics as key-value pairs, tags as run metadata, and artifacts, which are its output files. The run screen shows the run ID, the parameters, the metrics and run details, including a link to the source notebook. Files you uploaded with log_artifact appear in the Artifacts tab. Tags are key-value pairs you can use later to search for runs.
Logging is not unlimited. Since March 27, 2024, MLflow has enforced quotas on the total number of parameters, tags and metric steps per run, and on the number of runs per experiment. If you reach the runs-per-experiment quota, Databricks recommends deleting runs you no longer need. For the other quotas, it recommends changing your logging strategy to stay under them.
Checkpoint 7 of 7· Match them up
Match each piece of logged data to how the run stores it
Tap a term, then the definition that fits it.
Parameters and metrics are key-value pairs, tags are run metadata, and artifacts are output files stored with the run.
“model parameters and metrics saved as key-value pairs, tags for run metadata, and any artifacts, or output files, created by the run.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Because Databricks Autologging is on by default, a run you open with mlflow.start_run() automatically captures the model and parameters.Why is that wrong?
Databricks Autologging does not apply to runs created with mlflow.start_run(). You must call mlflow.autolog() (with exclusive=False to mix in manual logging).
Covered in Adding manual logging to an autologged run
2.With autologging enabled, evaluation scores on a test set and plots you create are logged for you.Why is that wrong?
Autologging records training metrics. A test-set score has to be logged with mlflow.log_metric, and a plot has to be saved to a file and logged with mlflow.log_artifact.
Covered in Logging metrics and artifacts by hand
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/mlflow/trackingOfficial docs
“The MLflow Tracking API logs parameters, metrics, tags, and artifacts from a model run.”
↩︎ Runs, experiments and the Tracking API“MLflow is pre-installed on Databricks Runtime ML clusters.”
↩︎ Runs, experiments and the Tracking API“to log additional artifacts such as CSV files or plots, use the MLflow logging API”
↩︎ Logging metrics and artifacts by hand“MLflow imposes a quota limit on the number of total parameters, tags, and metric steps for all existing and new runs”
↩︎ Where logged data shows up, and what limits apply“A run is a single execution of model code. During an MLflow run, you can log model parameters and results.”
↩︎ Key concept“to log additional artifacts such as CSV files or plots, use the MLflow logging API”
↩︎ Exam trap 2“If no active experiment is set, runs are logged to the notebook experiment.”
↩︎ Prediction“MLflow runs launched from any notebook or from the APIs can be logged to a workspace experiment.”
↩︎ Checkpoint - 2.
“When you use the mlflow.start_run() command in a notebook, the run logs metrics and parameters to the active experiment.”
↩︎ Runs, experiments and the Tracking API - 3.
“Log the ROC curve image file as an artifact”
↩︎ Logging metrics and artifacts by hand“The AUC score on test data is not automatically logged, so log it manually”
↩︎ Checkpoint - 4.
“In these cases, you must call mlflow.autolog() to save autologged content to the MLflow run.”
↩︎ Adding manual logging to an autologged run“Call mlflow.autolog() with exclusive=False.”
↩︎ Adding manual logging to an autologged run“Databricks Autologging is not applied to runs created using the MLflow fluent API with mlflow.start_run().”
↩︎ Exam trap 1“If you did not use with mlflow.start_run() in Step 2, end the MLflow run using mlflow.end_run().”
↩︎ Checkpoint - 5.https://docs.databricks.com/aws/en/mlflow/runsOfficial docs
“Artifacts saved from the run are available in the Artifacts tab.”
↩︎ Where logged data shows up, and what limits apply“The run screen shows the run ID, the parameters used for the run, the metrics resulting from the run, and details about the run”
↩︎ Where logged data shows up, and what limits apply“model parameters and metrics saved as key-value pairs, tags for run metadata, and any artifacts, or output files, created by the run.”
↩︎ Checkpoint