What you will be able to do
- Configure a GPU compute resource with Databricks Runtime ML and name the NVIDIA libraries it installs
- Identify the deep learning and distributed-training packages Databricks Runtime ML includes
- Predict which workloads Photon and Graviton speed up on the ML runtime, and which they don't
- Explain why Databricks Runtime ML uses Dedicated access mode for Unity Catalog, and when Databricks recommends AI Runtime instead
1.GPU support without installing drivers
Databricks Runtime for Machine Learning is the classic-compute runtime that comes with ML and deep learning libraries pre-installed, on CPU or GPU instances. On GPUs, it also saves you the work of setting up drivers and CUDA.
Creating GPU compute works like creating any other compute, with a few required settings. "The Machine learning checkbox must be checked. The GPU ML version is chosen automatically based on the worker type." The worker type must be a GPU instance type, and you can check Single node to get one GPU instance. With the Clusters API, you either set "use_ml_runtime": true (when kind = CLASSIC_PREVIEW) or set spark_version to a GPU-enabled version such as 15.4.x-gpu-ml-scala2.12.
On these instances, Databricks installs the NVIDIA driver and libraries on both the Spark driver and the workers:
- CUDA Toolkit, installed under /usr/local/cuda
- cuDNN: NVIDIA CUDA Deep Neural Network Library
- NCCL: NVIDIA Collective Communications Library
There is one exception. If you need a custom Docker image for GPU compute, you can't use the ML runtime. You must select a standard runtime version, and the custom GPU images are based on the official CUDA containers rather than on Databricks Runtime ML for GPU.
Checkpoint 1 of 7· Check yourself
You create GPU compute with Databricks Runtime ML. Which statement is correct?
Checking Machine learning and choosing a GPU worker type is enough. The GPU ML version follows from the worker type, and Databricks installs CUDA, cuDNN and NCCL for you.
“The Machine learning checkbox must be checked. The GPU ML version is chosen automatically based on the worker type.”Source: docs.databricks.com
Checkpoint 2 of 7· Exam question
A machine learning engineer is configuring a cluster to fine-tune a deep learning model and needs the cluster to use GPU-accelerated instance types with the correct NVIDIA CUDA drivers and GPU-enabled builds of TensorFlow and PyTorch already in place. What should the engineer do when creating the cluster?
Correct answer: B — Choose a GPU-enabled Databricks Runtime for Machine Learning version, preconfigured with CUDA drivers matched to that release and GPU-enabled TensorFlow and PyTorch builds.
- A. Compiling deep learning frameworks from source against a specific CUDA toolkit is exactly the manual, error-prone setup the ML runtime is designed to avoid, and doing it on a standard runtime forfeits the pre-built GPU stack entirely.
- B. The GPU variant of Databricks Runtime for Machine Learning is built and version-tested specifically for GPU instances, bundling the CUDA drivers and GPU-enabled TensorFlow and PyTorch builds so the engineer does not have to assemble that stack manually.
- C. GPU-enabled ML runtime images are tied to specific releases with matched driver and library versions; picking an arbitrary version without checking that it targets GPU hardware can leave the cluster without the correct CUDA stack for the instance type.
- D. Unity Catalog governs data and model access control and lineage; it has no role in detecting cluster hardware or installing GPU drivers and libraries at startup.
2.Deep learning frameworks and distributed training included
With the GPU stack in place, the frameworks are ready to use. PyTorch is included and "provides GPU accelerated tensor computation". Databricks Runtime ML also "includes TensorFlow and TensorBoard, so you can use these libraries without installing any packages."
Databricks recommends training neural networks on a single machine when possible, because distributed code is more complex and slower due to communication overhead. If the model or data is too large for one machine, Databricks Runtime ML already includes the packages you need to scale out:
- TorchDistributor: a PySpark module that lets you launch PyTorch training jobs as Spark jobs. - DeepSpeed distributor: built on top of TorchDistributor, for models that need more compute power but are limited by memory. - Ray: an open-source framework for parallel compute processing that scales ML workflows and AI applications.
The DeepSpeed distributor. It is built on TorchDistributor, and Databricks recommends it for models that need more compute power but are limited by memory constraints.
Checkpoint 3 of 7· Check yourself
Which packages does Databricks Runtime ML include for distributed training when a model or its data does not fit on one machine?
The runtime ships TorchDistributor, the DeepSpeed distributor and Ray for workloads too large for one machine.
“For these workloads, Databricks Runtime ML includes the TorchDistributor, DeepSpeed distributor and Ray packages.”Source: docs.databricks.com
3.Photon and Graviton: which workloads get faster
Starting with Databricks Runtime 15.2 ML, you can enable Photon on ML runtime compute. It helps the data side of an ML pipeline: it "improves performance for applications using Spark SQL, Spark DataFrames, feature engineering, GraphFrames, and xgboost4j." It does not help Spark RDDs, Pandas UDFs or non-JVM languages such as Python. That is why Python XGBoost gets no speedup while JVM-based xgboost4j does. Spark RDD APIs and Spark MLlib have limited compatibility with Photon and can run into Spark memory issues on large datasets. Photon is also not supported on GPU instance types, so the Photon checkbox must stay unchecked when you create GPU compute.
Graviton is the other option on CPU. Databricks Runtime 15.4 LTS ML and above support AWS Graviton instance types, which can improve performance for Spark, Photon, feature engineering, XGBoost, LightGBM and Spark MLlib gradient-boosting algorithms. Graviton instances may also offer better price-to-performance than other AWS EC2 instance types.
| Option | Minimum runtime | Helps | Does not help / not supported |
|---|---|---|---|
| Photon | Databricks Runtime 15.2 ML | Spark SQL, Spark DataFrames, feature engineering, GraphFrames, xgboost4j | Spark RDDs, Pandas UDFs, Python packages (XGBoost, PyTorch, TensorFlow); GPU instance types |
| Graviton instance types | Databricks Runtime 15.4 LTS ML | Spark, Photon, feature engineering, XGBoost, LightGBM, Spark MLlib gradient boosting | Not stated in the sources |
Checkpoint 4 of 7· Check yourself
Which workload on Databricks Runtime ML is most likely to speed up after you enable Photon?
Photon targets Spark SQL, DataFrames and feature engineering. Python frameworks, Pandas UDFs and RDDs don't benefit, and GPU instances don't support Photon.
“Photon improves performance for applications using Spark SQL, Spark DataFrames, feature engineering, GraphFrames, and xgboost4j.”Source: docs.databricks.com
Sources2
4.Dedicated access mode and Unity Catalog
Access to governed data comes with a requirement. "To access data in Unity Catalog on a compute resource running Databricks Runtime ML, you must set the access mode to Dedicated." You usually don't set this by hand: selecting the Machine learning checkbox in the create compute UI sets the access mode to Dedicated, with your account as the dedicated user.
Dedicated doesn't have to mean one person. You can assign the resource to a group in the Advanced section. The user's permissions then automatically down-scope to the group's permissions, so group members can share the compute securely. Runtime version also matters here: in dedicated access mode, fine-grained access control and querying tables created by Lakeflow pipelines (including streaming tables and materialized views) are available only on Databricks Runtime 15.4 LTS ML and above.
Checkpoint 5 of 7· Check yourself
A team wants to share one Databricks Runtime ML cluster and read Unity Catalog tables. What is the supported setup?
Unity Catalog access on the ML runtime requires Dedicated access mode, and assigning the resource to a group lets members share it securely through down-scoped permissions.
“the user's permissions automatically down-scope to the group's permissions”Source: docs.databricks.com
Sources2
5.Databricks Runtime ML versus AI Runtime for GPU work
The GPU support covered above is real, but Databricks now gives a more specific recommendation for some GPU work. For custom deep learning workloads on GPU compute, such as model fine-tuning or training, "Databricks recommends using AI Runtime instead of Databricks Runtime ML." AI Runtime is a serverless GPU compute offering (in Public Preview) with simplified setup, faster provisioning and better performance. It connects notebooks directly to serverless GPUs, offers A10 and H100 GPUs, and supports distributed training across multiple GPUs and nodes.
Databricks Runtime ML remains the classic-compute option for a comprehensive, ready-to-use environment covering both classic ML and deep learning, on CPU and GPU instances, with Photon, Graviton and Unity Catalog access through Dedicated mode.
| Aspect | Databricks Runtime ML | AI Runtime |
|---|---|---|
| Compute model | Classic compute | Serverless GPU compute (Public Preview) |
| Target workloads | Classic machine learning and deep learning | Custom single-node and multi-node deep learning, such as fine-tuning LLMs |
| Hardware | CPU and GPU instance types, including AWS Graviton | A10 GPUs and H100 GPUs |
| Environment | Pre-installed libraries such as PyTorch, TensorFlow and XGBoost | Default base environment, or an AI environment with packages like Transformers and Ray |
Checkpoint 6 of 7· Check yourself
A team will fine-tune a deep learning model on GPUs and wants fast provisioning without managing clusters. What does Databricks recommend?
For custom GPU deep learning such as fine-tuning, Databricks recommends AI Runtime over Databricks Runtime ML. Photon is not supported on GPUs, and Graviton instances are CPU instances.
“Databricks recommends using AI Runtime instead of Databricks Runtime ML”Source: docs.databricks.com
Checkpoint 7 of 7· Exam question
A team is migrating a PyTorch model training job that currently uses HorovodRunner to distribute training across multiple GPU workers on a Databricks Runtime for Machine Learning cluster. They want a supported path that does not depend on a library Databricks is phasing out of the ML runtime. Which approach should they take?
Correct answer: C — Rewrite the distributed training loop to use TorchDistributor, the API Databricks now recommends for distributed PyTorch training as Horovod is removed from newer ML runtime releases.
- A. Horovod and HorovodRunner are deprecated and are being removed from Databricks Runtime for Machine Learning in releases after the last LTS version that shipped them, so assuming indefinite support would leave the job broken on a future runtime upgrade.
- B. Lakeflow Spark Declarative Pipelines orchestrate data transformation workloads, not distributed deep learning training loops, so it is not a replacement for HorovodRunner in a PyTorch training job.
- C. TorchDistributor is the API Databricks now recommends for distributed PyTorch training now that Horovod and HorovodRunner are deprecated and are being removed from newer Databricks Runtime ML releases, making it the correct migration target.
- D. Databricks Runtime for Machine Learning continues to support distributed multi-GPU PyTorch training through TorchDistributor, so leaving the platform is unnecessary and discards the pre-configured GPU environment the runtime already provides.
Sources5
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Enabling Photon on an ML runtime GPU cluster will accelerate PyTorch or TensorFlow training.Why is that wrong?
Photon is not supported on GPU instance types, and Python packages such as PyTorch, TensorFlow and Python XGBoost get no improvement from it.
2.To use a custom Docker image on GPU compute, you start from Databricks Runtime ML for GPU.Why is that wrong?
Custom GPU images require a standard runtime version and are based on the official CUDA containers, not on Databricks Runtime ML for GPU.
Covered in GPU support without installing drivers
3.Databricks Runtime ML is still the recommended choice for every GPU deep learning workload.Why is that wrong?
For custom GPU deep learning such as fine-tuning or training, Databricks recommends AI Runtime, its serverless GPU offering.
Covered in Databricks Runtime ML versus AI Runtime for GPU work
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/compute/gpuOfficial docs
“Databricks installs the NVIDIA driver and libraries required to use GPUs on Spark driver and worker instances”
↩︎ GPU support without installing drivers“Photon is not supported with GPU instance types.”
↩︎ Exam trap 1“To create custom images for GPU compute, you must select a standard runtime version instead of Databricks Runtime ML for GPU.”
↩︎ Exam trap 2“The Machine learning checkbox must be checked. The GPU ML version is chosen automatically based on the worker type.”
↩︎ Checkpoint - 2.
“For GPU-based compute, select a GPU-enabled instance type in the Worker type drop-down menu.”
↩︎ GPU support without installing drivers“Graviton instances may also provide better price-to-performance value than other AWS EC2 instance types.”
↩︎ Photon and Graviton: which workloads get faster“To access data in Unity Catalog on a compute resource running Databricks Runtime ML, you must set the access mode to Dedicated.”
↩︎ Dedicated access mode and Unity Catalog“This automatically sets the access mode to Dedicated with your account as the dedicated user.”
↩︎ Dedicated access mode and Unity Catalog“AI Runtime is a serverless GPU compute offering optimized for deep learning with simplified setup, faster provisioning, and better performance.”
↩︎ Exam trap 3“Python packages such as XGBoost, PyTorch, and TensorFlow will not see an improvement with Photon.”
↩︎ Prediction“Photon improves performance for applications using Spark SQL, Spark DataFrames, feature engineering, GraphFrames, and xgboost4j.”
↩︎ Checkpoint“the user's permissions automatically down-scope to the group's permissions”
↩︎ Checkpoint“Databricks recommends using AI Runtime instead of Databricks Runtime ML”
↩︎ Checkpoint - 3.
“Databricks Runtime ML includes TensorFlow and TensorBoard, so you can use these libraries without installing any packages.”
↩︎ Deep learning frameworks and distributed training included - 4.
“When possible, Databricks recommends that you train neural networks on a single machine”
↩︎ Deep learning frameworks and distributed training included“For these workloads, Databricks Runtime ML includes the TorchDistributor, DeepSpeed distributor and Ray packages.”
↩︎ Checkpoint - 5.
“Serverless GPU compute environment optimized for custom single-node and multi-node deep learning workloads.”
↩︎ Databricks Runtime ML versus AI Runtime for GPU work