CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 1 · Lesson 12/48

    Log MLflow Models with log_model and Link Metrics to Them

    Manually log metrics, artifacts, and models in an MLflow Run.

    9 min read
    2.08% of exam
    4 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Log a trained model with mlflow.<model-type>.log_model and know which files MLflow adds automatically
    • Explain how an MLflow 3 LoggedModel differs from a model stored as a run artifact, and link metrics to it with model_id
    • Capture extra dependencies with extra_pip_requirements, code_paths and the artifacts argument
    • Load a logged model back from a runs:/ path

    1.What log_model writes

    Logging a metric stores a number and logging an artifact stores a file. Logging a model stores something more structured. An MLflow Model is a standard packaging format that downstream tools understand, from Spark batch inference to REST serving. The format lets you save a model in different flavors, such as python-function, pytorch or sklearn. That is why the call is namespaced by flavor: mlflow.<model-type>.log_model(model, ...), for example mlflow.sklearn.log_model or mlflow.pytorch.log_model.

    These environment files let you recreate the training environment with virtualenv (recommended) or conda. For built-in flavors MLflow can record dependencies reliably. For example, mlflow.sklearn.log_model logs the scikit-learn version. Autologging with those libraries does the same. When you train inside a run without autologging, log_model is the manual call that puts the model into the run.

    Checkpoint 1 of 5· Exam question

    A model evaluation script on Databricks saves a confusion matrix as `confusion_matrix.png` to local disk and needs it attached to the current MLflow run under an artifact subfolder named `plots`, so it appears grouped separately from other run files in the UI. Which call correctly does this?

    Sources12

    2.MLflow 3: the model as its own object

    In MLflow 2.x, a logged model is stored as a run artifact and appears in the run's Artifacts tab. MLflow 3 adds a dedicated LoggedModel object, which log_model() creates with a unique ID. A LoggedModel carries its own metadata, such as parameters and metrics. It persists across the model's lifecycle and links to artifacts, metrics, parameters and the generating code. In deep learning training, MLflow creates a separate LoggedModel for each checkpoint, so you can compare checkpoints. The logging call stays the same log_model() API. In MLflow 3 you name the model with name= and can pass params= directly:

    Logging a scikit-learn model as an MLflow 3 LoggedModel and inspecting itpython
    ## Log the model
    model_info = mlflow.sklearn.log_model(
      sk_model=lr,
      name="elasticnet",
      params={
        "alpha": 0.5,
        "l1_ratio": 0.5,
      },
      input_example = train_x
    )
    
    # Inspect the LoggedModel and its properties
    logged_model = mlflow.get_logged_model(model_info.model_id)
    print(logged_model.model_id, logged_model.params)

    log_model returns a model_info with a model_id. Pass that ID to mlflow.log_metric or mlflow.log_metrics, and the evaluation metrics attach to the LoggedModel as well as to the run. You can also pass the dataset the metrics were computed on. To see the results, open the experiment and select the Models tab. It lists every LoggedModel in the experiment with its metrics, parameters and artifacts.

    Checkpoint 2 of 5· Fill the gap

    Which parameter links these metrics to the LoggedModel created above?

    mlflow.log_metrics(
      metrics={
        "rmse": rmse,
        "r2": r2,
        "mae": mae,
      },
       ? =logged_model.model_id,
      dataset=train_dataset
    )

    Checkpoint 3 of 5· Check yourself

    In MLflow 3, where do you find all models logged in an experiment, together with their metrics and parameters?

    Sources34

    3.Logging what MLflow cannot infer

    Some libraries can be installed with pip but have no built-in flavor. For those, use mlflow.pyfunc.log_model. MLflow infers dependencies with mlflow.models.infer_pip_requirements and writes them to requirements.txt as a model artifact. Log requirements with the exact version, for example f"nltk=={nltk.__version__}" instead of just nltk. When inference misses something, there are several ways to add to or override the model's environment:

    Ways to attach dependencies and extra files when logging a model
    OptionWhat it doesGuidance
    extra_pip_requirementsAdds requirements on top of the ones MLflow infersUse when automatic inference misses a library
    pip_requirements / conda_envReplaces the whole set of requirementsGenerally discouraged because it overrides what MLflow picks up automatically
    code_paths (code_path in MLflow 2.x)Stores .py files, directories or wheels in a code directory with the model and adds them to the Python path at loadFor code you cannot install with %pip install
    artifactsLogs an extra artifact, such as a .txt or .json list of non-Python dependenciesRecommended because MLflow does not detect Java, R or native packages
    MLFLOW_LOCK_MODEL_DEPENDENCIESEnvironment variable that makes MLflow 3 capture direct and transitive dependenciesSet to "true" before logging
    MLflow 3 pyfunc logging with local code files attachedpython
    mlflow.pyfunc.log_model(
       name=name,
       code_paths=[filename.py],
       data_path=data_path,
       conda_env=conda_env,
    )

    Checkpoint 4 of 5· Check yourself

    Your model calls a Linux system package at inference time. What does Databricks recommend so the serving environment knows about it?

    Sources2

    4.Loading the model back from the run

    You can check that a model was logged by loading it again. mlflow.<model-type>.load_model(modelpath) accepts several path forms. Examples are a run-relative path runs:/{run_id}/{model-path}, an MLflow 3 models:/{model_id} path and a Unity Catalog volumes path. For Python models, mlflow.pyfunc.load_model() loads any flavor as a generic Python function. You can also turn the model into a Spark UDF with mlflow.pyfunc.spark_udf for batch or streaming scoring.

    Loading the model logged by a run through its runs:/ URIpython
    model_loaded = mlflow.pyfunc.load_model(
      'runs:/{run_id}/model'.format(
        run_id=run.info.run_id
      )
    )

    You don't have to write this code by hand. When you log a model in a Databricks notebook, Databricks generates ready-to-copy snippets for loading it and predicting on Spark or pandas DataFrames.

    Checkpoint 5 of 5· Put it in order

    Put the steps for finding the generated load-and-predict snippet in order

    1. 1.Click the name of the logged model to open the code panel
    2. 2.Navigate to the Runs screen for the run that generated the model
    3. 3.Scroll to the Artifacts section

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.log_model captures every dependency, including Java, R and native system packages.Why is that wrong?

      MLflow infers only Python dependencies. You have to log non-Python dependencies yourself, for example as a dependency-list artifact.

      Covered in Logging what MLflow cannot infer

    2. 2.If inference misses one library, overwrite the requirements with pip_requirements.Why is that wrong?

      Add the missing library with extra_pip_requirements. Replacing the full set with pip_requirements or conda_env is discouraged because it throws away the dependencies MLflow inferred.

      Covered in Logging what MLflow cannot infer

    Practise it for real

    In a Databricks notebook, train a scikit-learn classifier, manually log a test metric and a plot artifact, then load the logged model back from its run.

    1. 1.Call mlflow.sklearn.autolog(), then train a GradientBoostingClassifier inside with mlflow.start_run(run_name='gradient_boost') as run:

      Why: Autologging records the model, parameters and training metrics for a run you opened yourself.

      You should see: A new run named gradient_boost appears in the notebook's Experiment Runs sidebar.

    2. 2.Inside the same run, compute roc_auc_score on the test set and call mlflow.log_metric("test_auc", roc_auc).

      Why: Test-set scores are not autologged.

      You should see: test_auc is listed among the run's metrics.

    3. 3.Save the ROC curve with roc_curve.figure_.savefig("roc_curve.png") and call mlflow.log_artifact("roc_curve.png").

      Why: log_artifact uploads a file from local disk, so the plot has to be written to a file first.

      You should see: roc_curve.png is in the run's Artifacts tab.

    4. 4.After the run, load the model with mlflow.pyfunc.load_model('runs:/{run_id}/model'.format(run_id=run.info.run_id)) and compare its predictions with model.predict(X_test).

      Why: This shows the logged model can be loaded from the run alone.

      You should see: np.array_equal on the two prediction arrays returns True.

    Stuck? Get a nudge

    If the run is missing from the sidebar, click the refresh icon in the experiment sidebar.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “To log a model to the MLflow tracking server, use mlflow.<model-type>.log_model(model, ...).”
      ↩︎ What log_model writes
      “The format defines a convention that lets you save a model in different flavors (python-function, pytorch, sklearn, and so on)”
      ↩︎ What log_model writes
      “a run-relative path (such as runs:/{run_id}/{model-path})”
      ↩︎ Loading the model back from the run
      “For Python MLflow models, an additional option is to use mlflow.pyfunc.load_model() to load the model as a generic Python function.”
      ↩︎ Loading the model back from the run
      “When you log a model, MLflow automatically logs requirements.txt and conda.yaml files.”
      ↩︎ Prediction
      “Databricks automatically generates code snippets that you can copy and use to load and run the model”
      ↩︎ Checkpoint
    2. 2.
      “MLflow supports scikit-learn in the mlflow.sklearn module, and the command mlflow.sklearn.log_model logs the sklearn version.”
      ↩︎ What log_model writes
      “you can specify additional dependencies with the extra_pip_requirements parameter in the log_model command”
      ↩︎ Logging what MLflow cannot infer
      “MLflow stores any files or directories passed using code_paths or code_path as artifacts along with the model in a code directory.”
      ↩︎ Logging what MLflow cannot infer
      “Be sure to log the requirements with the exact library version”
      ↩︎ Logging what MLflow cannot infer
      “MLflow does not automatically pick up non-Python dependencies, such as Java packages, R packages, and native packages”
      ↩︎ Exam trap 1
      “but doing so is generally discouraged because this overrides the dependencies which MLflow picks up automatically”
      ↩︎ Exam trap 2
      “mlflow.pyfunc.log_model allows you to specify this additional artifact using the artifacts argument.”
      ↩︎ Checkpoint
    3. 3.
      “When you train a model, use mlflow.<model-flavor>.log_model() to create a LoggedModel that ties together all of its critical information using a unique ID.”
      ↩︎ MLflow 3: the model as its own object
      “These metrics are now linked to the LoggedModel entity”
      ↩︎ MLflow 3: the model as its own object
    4. 4.
      “If you log a model from a run, the model appears in the Artifacts tab”
      ↩︎ MLflow 3: the model as its own object
      “In MLflow 3, models are now their distinct first-class object rather than being logged as run artifact.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise Databricks Certified Machine Learning Associate in quiz mode.

    Spotted a mistake, or was something unclear? Tell us.