What you will be able to do
- Log a trained model with mlflow.<model-type>.log_model and know which files MLflow adds automatically
- Explain how an MLflow 3 LoggedModel differs from a model stored as a run artifact, and link metrics to it with model_id
- Capture extra dependencies with extra_pip_requirements, code_paths and the artifacts argument
- Load a logged model back from a runs:/ path
1.What log_model writes
Logging a metric stores a number and logging an artifact stores a file. Logging a model stores something more structured. An MLflow Model is a standard packaging format that downstream tools understand, from Spark batch inference to REST serving. The format lets you save a model in different flavors, such as python-function, pytorch or sklearn. That is why the call is namespaced by flavor: mlflow.<model-type>.log_model(model, ...), for example mlflow.sklearn.log_model or mlflow.pytorch.log_model.
These environment files let you recreate the training environment with virtualenv (recommended) or conda. For built-in flavors MLflow can record dependencies reliably. For example, mlflow.sklearn.log_model logs the scikit-learn version. Autologging with those libraries does the same. When you train inside a run without autologging, log_model is the manual call that puts the model into the run.
Checkpoint 1 of 5· Exam question
A model evaluation script on Databricks saves a confusion matrix as `confusion_matrix.png` to local disk and needs it attached to the current MLflow run under an artifact subfolder named `plots`, so it appears grouped separately from other run files in the UI. Which call correctly does this?
Correct answer: A — `mlflow.log_artifact("confusion_matrix.png", artifact_path="plots")`, uploading the single local file into the `plots` subdirectory of the run's artifact store.
- A. This is correct because `log_artifact` uploads exactly one local file, and its `artifact_path` argument places that file under the named subdirectory within the run's artifact store.
- B. This is incorrect because the roles are reversed: `log_artifacts` expects a local directory of files, while `log_artifact` is the one that accepts a single file path.
- C. This is incorrect because `log_metric` requires a numeric value and an integer `step`, so passing a file path string and a folder name as `step` does not match the expected types.
- D. This is incorrect because `log_model` expects a fitted model object paired with a supported flavor module, not an arbitrary image file, so it cannot be used to store a PNG.
2.MLflow 3: the model as its own object
In MLflow 2.x, a logged model is stored as a run artifact and appears in the run's Artifacts tab. MLflow 3 adds a dedicated LoggedModel object, which log_model() creates with a unique ID. A LoggedModel carries its own metadata, such as parameters and metrics. It persists across the model's lifecycle and links to artifacts, metrics, parameters and the generating code. In deep learning training, MLflow creates a separate LoggedModel for each checkpoint, so you can compare checkpoints. The logging call stays the same log_model() API. In MLflow 3 you name the model with name= and can pass params= directly:
## Log the model
model_info = mlflow.sklearn.log_model(
sk_model=lr,
name="elasticnet",
params={
"alpha": 0.5,
"l1_ratio": 0.5,
},
input_example = train_x
)
# Inspect the LoggedModel and its properties
logged_model = mlflow.get_logged_model(model_info.model_id)
print(logged_model.model_id, logged_model.params)log_model returns a model_info with a model_id. Pass that ID to mlflow.log_metric or mlflow.log_metrics, and the evaluation metrics attach to the LoggedModel as well as to the run. You can also pass the dataset the metrics were computed on. To see the results, open the experiment and select the Models tab. It lists every LoggedModel in the experiment with its metrics, parameters and artifacts.
Checkpoint 2 of 5· Fill the gap
Which parameter links these metrics to the LoggedModel created above?
mlflow.log_metrics(
metrics={
"rmse": rmse,
"r2": r2,
"mae": mae,
},
? =logged_model.model_id,
dataset=train_dataset
)Passing model_id ties the logged metrics to the LoggedModel entity. Without it they belong only to the active run.
Source: docs.databricks.comCheckpoint 3 of 5· Check yourself
In MLflow 3, where do you find all models logged in an experiment, together with their metrics and parameters?
MLflow 3 makes models first-class objects. The experiment's Models tab lists every LoggedModel with its metrics, parameters and artifacts.
“In MLflow 3, models are now their distinct first-class object rather than being logged as run artifact.”Source: docs.databricks.com
3.Logging what MLflow cannot infer
Some libraries can be installed with pip but have no built-in flavor. For those, use mlflow.pyfunc.log_model. MLflow infers dependencies with mlflow.models.infer_pip_requirements and writes them to requirements.txt as a model artifact. Log requirements with the exact version, for example f"nltk=={nltk.__version__}" instead of just nltk. When inference misses something, there are several ways to add to or override the model's environment:
| Option | What it does | Guidance |
|---|---|---|
| extra_pip_requirements | Adds requirements on top of the ones MLflow infers | Use when automatic inference misses a library |
| pip_requirements / conda_env | Replaces the whole set of requirements | Generally discouraged because it overrides what MLflow picks up automatically |
| code_paths (code_path in MLflow 2.x) | Stores .py files, directories or wheels in a code directory with the model and adds them to the Python path at load | For code you cannot install with %pip install |
| artifacts | Logs an extra artifact, such as a .txt or .json list of non-Python dependencies | Recommended because MLflow does not detect Java, R or native packages |
| MLFLOW_LOCK_MODEL_DEPENDENCIES | Environment variable that makes MLflow 3 capture direct and transitive dependencies | Set to "true" before logging |
mlflow.pyfunc.log_model(
name=name,
code_paths=[filename.py],
data_path=data_path,
conda_env=conda_env,
)Checkpoint 4 of 5· Check yourself
Your model calls a Linux system package at inference time. What does Databricks recommend so the serving environment knows about it?
MLflow only infers Python dependencies. For non-Python packages, Databricks recommends logging a dependency list as an artifact with the model.
“mlflow.pyfunc.log_model allows you to specify this additional artifact using the artifacts argument.”Source: docs.databricks.com
Sources2
4.Loading the model back from the run
You can check that a model was logged by loading it again. mlflow.<model-type>.load_model(modelpath) accepts several path forms. Examples are a run-relative path runs:/{run_id}/{model-path}, an MLflow 3 models:/{model_id} path and a Unity Catalog volumes path. For Python models, mlflow.pyfunc.load_model() loads any flavor as a generic Python function. You can also turn the model into a Spark UDF with mlflow.pyfunc.spark_udf for batch or streaming scoring.
model_loaded = mlflow.pyfunc.load_model(
'runs:/{run_id}/model'.format(
run_id=run.info.run_id
)
)You don't have to write this code by hand. When you log a model in a Databricks notebook, Databricks generates ready-to-copy snippets for loading it and predicting on Spark or pandas DataFrames.
Checkpoint 5 of 5· Put it in order
Put the steps for finding the generated load-and-predict snippet in order
- 1.Click the name of the logged model to open the code panel
- 2.Navigate to the Runs screen for the run that generated the model
- 3.Scroll to the Artifacts section
The snippets are attached to the logged model inside the run's Artifacts section. Clicking the model name opens a side panel with the code.
“Databricks automatically generates code snippets that you can copy and use to load and run the model”Source: docs.databricks.com
It shows that what was logged is a faithful copy of the trained model. Another notebook or job could load it through the runs:/ path and get the same predictions.
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.log_model captures every dependency, including Java, R and native system packages.Why is that wrong?
MLflow infers only Python dependencies. You have to log non-Python dependencies yourself, for example as a dependency-list artifact.
Covered in Logging what MLflow cannot infer
2.If inference misses one library, overwrite the requirements with pip_requirements.Why is that wrong?
Add the missing library with extra_pip_requirements. Replacing the full set with pip_requirements or conda_env is discouraged because it throws away the dependencies MLflow inferred.
Covered in Logging what MLflow cannot infer
Practise it for real
In a Databricks notebook, train a scikit-learn classifier, manually log a test metric and a plot artifact, then load the logged model back from its run.
1.Call mlflow.sklearn.autolog(), then train a GradientBoostingClassifier inside with mlflow.start_run(run_name='gradient_boost') as run:
Why: Autologging records the model, parameters and training metrics for a run you opened yourself.
You should see: A new run named gradient_boost appears in the notebook's Experiment Runs sidebar.
2.Inside the same run, compute roc_auc_score on the test set and call mlflow.log_metric("test_auc", roc_auc).
Why: Test-set scores are not autologged.
You should see: test_auc is listed among the run's metrics.
3.Save the ROC curve with roc_curve.figure_.savefig("roc_curve.png") and call mlflow.log_artifact("roc_curve.png").
Why: log_artifact uploads a file from local disk, so the plot has to be written to a file first.
You should see: roc_curve.png is in the run's Artifacts tab.
4.After the run, load the model with mlflow.pyfunc.load_model('runs:/{run_id}/model'.format(run_id=run.info.run_id)) and compare its predictions with model.predict(X_test).
Why: This shows the logged model can be loaded from the run alone.
You should see: np.array_equal on the two prediction arrays returns True.
Stuck? Get a nudge
If the run is missing from the sidebar, click the refresh icon in the experiment sidebar.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/mlflow/modelsOfficial docs
“To log a model to the MLflow tracking server, use mlflow.<model-type>.log_model(model, ...).”
↩︎ What log_model writes“The format defines a convention that lets you save a model in different flavors (python-function, pytorch, sklearn, and so on)”
↩︎ What log_model writes“a run-relative path (such as runs:/{run_id}/{model-path})”
↩︎ Loading the model back from the run“For Python MLflow models, an additional option is to use mlflow.pyfunc.load_model() to load the model as a generic Python function.”
↩︎ Loading the model back from the run“When you log a model, MLflow automatically logs requirements.txt and conda.yaml files.”
↩︎ Prediction“Databricks automatically generates code snippets that you can copy and use to load and run the model”
↩︎ Checkpoint - 2.
“MLflow supports scikit-learn in the mlflow.sklearn module, and the command mlflow.sklearn.log_model logs the sklearn version.”
↩︎ What log_model writes“you can specify additional dependencies with the extra_pip_requirements parameter in the log_model command”
↩︎ Logging what MLflow cannot infer“MLflow stores any files or directories passed using code_paths or code_path as artifacts along with the model in a code directory.”
↩︎ Logging what MLflow cannot infer“Be sure to log the requirements with the exact library version”
↩︎ Logging what MLflow cannot infer“MLflow does not automatically pick up non-Python dependencies, such as Java packages, R packages, and native packages”
↩︎ Exam trap 1“but doing so is generally discouraged because this overrides the dependencies which MLflow picks up automatically”
↩︎ Exam trap 2“mlflow.pyfunc.log_model allows you to specify this additional artifact using the artifacts argument.”
↩︎ Checkpoint - 3.
“When you train a model, use mlflow.<model-flavor>.log_model() to create a LoggedModel that ties together all of its critical information using a unique ID.”
↩︎ MLflow 3: the model as its own object“These metrics are now linked to the LoggedModel entity”
↩︎ MLflow 3: the model as its own object - 4.https://docs.databricks.com/aws/en/mlflow/runsOfficial docs
“If you log a model from a run, the model appears in the Artifacts tab”
↩︎ MLflow 3: the model as its own object“In MLflow 3, models are now their distinct first-class object rather than being logged as run artifact.”
↩︎ Checkpoint