Subdomain 1.2: Identify the advantages of using ML runtimes
1.A team maintains a model training pipeline that must produce identical results when re-run six months later, including the exact versions of scikit-learn, XGBoost, and their transitive dependencies used at training time. Which characteristic of Databricks Runtime for Machine Learning most directly supports this requirement?
- A.Databricks Runtime for Machine Learning automatically upgrades scikit-learn and XGBoost to their latest release on every cluster restart, keeping the environment current with upstream fixes.
- B.Each Databricks Runtime for Machine Learning release pins a tested, documented set of library versions, so re-running a job on the same runtime version reproduces the same dependency set.
- C.Databricks Runtime for Machine Learning stores every notebook's output as a Delta table, which independently guarantees identical library versions across separate training runs.
- D.Databricks Runtime for Machine Learning removes the need to track library versions at all, because MLflow autologging reconstructs the original environment from the model artifact alone.
Show answer & explanation
Correct answer: B — Each Databricks Runtime for Machine Learning release pins a tested, documented set of library versions, so re-running a job on the same runtime version reproduces the same dependency set.
- A. Databricks Runtime for Machine Learning does not silently auto-upgrade libraries on restart; a given runtime release keeps a fixed library set, which is what makes it reproducible rather than a moving target.
- B. Each Databricks Runtime for Machine Learning release ships a specific, tested, and documented set of library versions, so pinning a job to that runtime version reproduces the same scikit-learn, XGBoost, and dependency versions on a later run.
- C. Saving notebook output to Delta tables preserves data, not the software environment; it does not record or guarantee which library versions produced that output on a separate run.
- D. MLflow autologging can record library versions as part of a model's metadata, but it does not eliminate the need to track versions during training or reconstruct an environment on its own without a pinned runtime to install against.