CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 1 · Lesson 2/48

    Databricks Runtime ML: Pre-Installed Libraries and When to Use It

    Identify the advantages of using ML runtimes

    9 min read
    2.08% of exam
    5 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what Databricks Runtime for Machine Learning provides compared with building an ML environment yourself
    • Identify the top-tier libraries and how Databricks updates, deprecates and removes pre-installed libraries
    • Choose the right way to add extra libraries to a Databricks Runtime ML environment
    • Decide when a workload justifies selecting the machine learning runtime and when it does not

    Key concept

    Databricks Runtime for Machine Learning (Databricks Runtime ML) — A classic-compute runtime that sets up a compute resource with machine learning and deep learning infrastructure already in place, including the common ML and DL libraries. You skip installing and reconciling a library stack and start from an environment Databricks has already put together and tested.

    1.What the ML runtime gives you out of the box

    On classic compute, the main advantage of Databricks Runtime ML is that you don't build the environment yourself. Without it, you install PyTorch, scikit-learn, XGBoost, MLflow and the rest one by one, then work out which versions run together. With it, those libraries come pre-installed. Databricks describes the result as pre-configured cluster environments with major ML libraries "pre-installed and tested together", and these are available for both CPU and GPU-accelerated clusters.

    The training documentation lists pre-installed libraries as the first key feature. It names PyTorch, TensorFlow and XGBoost and says they "receive frequent updates and optimized support." The runtime is pitched at users who want one ready-to-use environment for classic machine learning and for deep learning, so a single cluster can run a gradient-boosted model in one notebook and a neural network in another.

    The environment also moves forward over time. Each Databricks Runtime ML release refreshes the bundled libraries, so picking a newer runtime version also gets you newer library versions, without upgrading packages yourself.

    Checkpoint 1 of 4· Exam question

    A data science team currently spins up a general-purpose Databricks Runtime cluster and then runs a setup script on every cluster start to install scikit-learn, XGBoost, and MLflow before any notebook can execute. The team wants to remove this manual step and standardize the environment for every new project. Which change addresses this directly?

    Sources123

    2.Top-tier libraries and the maintenance policy

    Not every pre-installed library is treated the same way. Databricks designates a subset as top-tier libraries, and for these "Databricks provides a faster update cadence, updating to the latest package releases with each runtime release (barring dependency conflicts)." Top-tier libraries also get advanced support, testing and embedded optimizations. The list changes only at major releases.

    The full top-tier list is: datasets, GraphFrames, MLflow, PyTorch, Scikit-learn, streaming, TensorBoard and transformers. The list leaves some out. XGBoost is pre-installed but not top-tier. TensorFlow is still included, but starting with Databricks Runtime 18.0 ML, TensorFlow and spark-tensorflow-connector are no longer top-tier.

    The maintenance policy has two separate thresholds that are easy to confuse:

    - Dropped from the top-tier list: possible if the library has no new commits in two months and no new releases in more than six months, if its usage drops significantly, or if it is replaced by new packages. It can return to the list when active maintenance resumes. - Removed from the runtime entirely: happens when the library is no longer actively maintained (for example, no new commits in three months and no new releases in more than nine months, an archived repository, or an announced stop in maintenance), or when no stable release works on the new runtime.

    Before a removal, Databricks warns you in three places: the runtime release notes, a notification when you import the library, and the documentation. The notice says the library will be removed in the next major Databricks Runtime ML release. If you still need a removed library, you can install it manually or stay on an earlier runtime version.

    Two different maintenance thresholds in Databricks Runtime ML
    OutcomeInactivity triggerOther triggers
    Removed from the top-tier list (still installed)No new commits in two months and no new releases in more than six monthsUsage drops significantly; replaced by new packages that fill major gaps
    Removed from the runtimeNo new commits in three months and no new releases in more than nine monthsRepository archived; announced stop in maintenance; no stable release functional for the new runtime

    Checkpoint 2 of 4· Match them up

    Match each maintenance-policy situation to what Databricks does

    Tap a term, then the definition that fits it.

    Sources4

    3.Adding libraries beyond the pre-installed set

    A pre-built environment can still be extended. Databricks says you can install additional libraries to create a custom environment for your notebook or compute resource, and the right method depends on who needs the library.

    To make a library available to every notebook on a compute resource, "create a compute-scoped library". You can also install libraries with an init script when the compute is created. If only one notebook session needs the library, use notebook-scoped Python libraries.

    Databricks' compute best practices prefer the narrower option. Installing libraries at the compute level "creates environment drift across jobs," and init scripts can introduce library conflicts that make environments less predictable. The recommended alternative is to "use %pip install in notebooks or define dependencies in an environment spec." This also makes classic workloads easier to move to serverless later.

    Ways to add a library on top of Databricks Runtime ML
    MethodScopeBest-practice note
    Compute-scoped libraryAll notebooks running on the compute resourceCreates environment drift across jobs
    Init scriptInstalled during compute creationCan introduce library conflicts and less predictable environments
    Notebook-scoped Python libraries (%pip install)A specific notebook session onlyRecommended, along with an environment spec

    Checkpoint 3 of 4· Check yourself

    A data scientist needs an extra package for one experiment notebook on a shared Databricks Runtime ML cluster. Which approach matches Databricks guidance?

    Sources35

    4.When the ML runtime is the right choice

    The feature that makes the ML runtime useful can also cause problems. Because Databricks Runtime ML "installs a large set of libraries that can conflict with your own dependencies if not needed, causing errors or silent correctness issues," Databricks' compute best practices name three workloads that justify it: GPUs, distributed ML training and AutoML.

    For those workloads, the bundled libraries are what you need: deep learning frameworks, GPU-ready builds and distributed-training packages. For a workload outside those three, a standard runtime plus a few notebook-scoped installs gives you a smaller environment with fewer chances of a version conflict you didn't notice.

    Checkpoint 4 of 4· Check yourself

    Which risk does Databricks name for selecting the machine learning runtime when a workload does not need it?

    Sources5

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Every library pre-installed in Databricks Runtime ML, including XGBoost and TensorFlow, is a top-tier library with the fastest update cadence.Why is that wrong?

      Top-tier is a named subset: datasets, GraphFrames, MLflow, PyTorch, Scikit-learn, streaming, TensorBoard and transformers. XGBoost is not on it, and TensorFlow was dropped at 18.0 ML.

      Covered in Top-tier libraries and the maintenance policy

    2. 2.Losing top-tier status means a library is removed from the runtime.Why is that wrong?

      These are separate outcomes with separate thresholds. Two months without commits and six months without releases can drop a library from the top-tier list. Removal from the runtime needs three months and nine months, or another sign that maintenance has ended.

      Covered in Top-tier libraries and the maintenance policy

    3. 3.The ML runtime is always the safer default for any data science work because it has more libraries.Why is that wrong?

      Databricks advises selecting it only for GPU, distributed ML training or AutoML workloads, because the large library set can conflict with your own dependencies.

      Covered in When the ML runtime is the right choice

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “pre-configured cluster environments with major ML libraries pre-installed and tested together, for both CPU and GPU-accelerated clusters”
      ↩︎ What the ML runtime gives you out of the box
    2. 2.
      “Includes popular libraries like PyTorch, TensorFlow, and XGBoost, which receive frequent updates and optimized support.”
      ↩︎ What the ML runtime gives you out of the box
    3. 3.
      “The libraries are updated with each release to include new features and fixes.”
      ↩︎ What the ML runtime gives you out of the box
      “To make a library available for all notebooks running on a compute resource, create a compute-scoped library.”
      ↩︎ Adding libraries beyond the pre-installed set
      “automates the creation of a compute resource with pre-built machine learning and deep learning infrastructure including the most common ML and DL libraries”
      ↩︎ Key concept
    4. 4.
      “Databricks provides a faster update cadence, updating to the latest package releases with each runtime release (barring dependency conflicts).”
      ↩︎ Top-tier libraries and the maintenance policy
      “you can either install the library manually or use an earlier version of Databricks Runtime ML”
      ↩︎ Top-tier libraries and the maintenance policy
      “Starting with Databricks Runtime 18.0 ML, TensorFlow and spark-tensorflow-connector are no longer top-tier libraries.”
      ↩︎ Exam trap 1
      “If the library has no new commits in two months and no new releases in more than six months.”
      ↩︎ Exam trap 2
      “Databricks also provides advanced support, testing, and embedded optimizations for top-tier libraries.”
      ↩︎ Prediction
      “No new commits in three months and no new releases in more than nine months.”
      ↩︎ Checkpoint
    5. 5.
      “Installing libraries at the compute level creates environment drift across jobs.”
      ↩︎ Adding libraries beyond the pre-installed set
      “Only select a machine learning runtime if your workload uses GPUs, distributed ML training, or AutoML.”
      ↩︎ When the ML runtime is the right choice
      “Only select a machine learning runtime if your workload uses GPUs, distributed ML training, or AutoML.”
      ↩︎ Exam trap 3
      “Instead, use %pip install in notebooks or define dependencies in an environment spec.”
      ↩︎ Checkpoint
      “installs a large set of libraries that can conflict with your own dependencies if not needed, causing errors or silent correctness issues”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Databricks Runtime ML on GPUs, Photon, Graviton and Unity Catalog

    Spotted a mistake, or was something unclear? Tell us.