CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 31/48

    Building a Model Training Pipeline on Databricks: Structure, scikit-learn Pipelines, and Metrics

    Develop a training pipeline

    7 min read
    2.08% of exam
    4 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Name the tasks inside a Databricks model training pipeline and say what each one logs to MLflow
    • Explain the scikit-learn contract that a Pipeline enforces between transformers and the final estimator
    • Pick a primary metric and a class-weighting strategy using the AutoML classification and regression parameters

    Key concept

    Model training pipeline — Code that trains and tunes a model, then evaluates it on held-out data. Every parameter, metric and artifact is logged to MLflow, so the final model can be traced back to the data and code that produced it.

    1.What a training pipeline contains

    In the Databricks MLOps workflow, data scientists write the model training pipeline in the development environment, reading from tables in the dev or prod catalogs. The pipeline has two tasks. The first is training and tuning. Model parameters, metrics and artifacts are logged to the MLflow Tracking server, and after hyperparameter tuning the final model artifact is logged too, which records the link between the model, its input data and its code. The second is evaluation. The model is tested on held-out data, and the results are logged as well. Evaluation answers one question: does the new model beat the current production model?

    In MLflow terms, a *run* is one execution of model code, and an *experiment* is a collection of related runs. Inside an experiment you can compare runs to see how performance depends on parameter settings. If you set no experiment, runs from a notebook go to the notebook experiment. To log to a shared workspace experiment instead, set it explicitly:

    Logging runs to a workspace experiment instead of the notebook experimentpython
    experiment_name = "/Shared/name_of_experiment/"
    mlflow.set_experiment(experiment_name)

    Reproducibility also needs care. Many algorithms have random elements, and many libraries let you set a seed to fix the starting conditions for them. Seeds don't control everything, though. Some algorithms are sensitive to data order, and distributed algorithms can be affected by how the data is partitioned. Databricks recommends the PySpark functions repartition and sortWithinPartitions to control that variation. Once training finishes, register the model to Unity Catalog in the catalog of the environment where the pipeline ran. The training task produces a model URI, which the next validation task can use.

    Checkpoint 1 of 4· Check yourself

    A newly trained model has been logged. What is the evaluation task in the training pipeline meant to decide?

    Sources12

    2.Implementing the training step with a scikit-learn Pipeline

    "Pipeline" has two meanings in this lesson. The *training pipeline* above is the whole workflow task. Inside it, the training code often uses scikit-learn's Pipeline class, which chains preprocessing and modelling steps into a single estimator. Databricks provides an example notebook that uses MLflow autologging with scikit-learn, so a fitted Pipeline can be tracked without extra logging code.

    scikit-learn separates transformers, which reshape data with fit and transform, from predictors such as classifiers and regressors, which fit and then predict. A Pipeline enforces one rule: every step except the last must be a transformer. The last step can be a transformer, a predictor or a clustering estimator. The pipeline then exposes whatever methods the last step has. If the last step can predict, calling predict on the pipeline runs the data through all the earlier transforms and passes the result to that final step.

    Pipelines also connect to target transformations. scikit-learn's TransformedTargetRegressor handles transforming the target, for example a log-transform of y. The available sources say nothing more about converting predictions or metrics back to the original scale after a log transform. Check the scikit-learn docs for that detail.

    Checkpoint 2 of 4· Check yourself

    You build a Pipeline with steps [StandardScaler, LogisticRegression, PCA]. Why is this invalid?

    Sources23

    3.Choosing a metric and weighting imbalanced classes

    The pipeline needs a metric before evaluation means anything. The Databricks AutoML API lists the metrics a run can rank models by, which also makes it a handy checklist of the standard options. Classification can use F1 (the default), log loss, precision, accuracy and ROC AUC. Regression can use R-squared (the default), MAE, RMSE and MSE. For binary classification, pos_label sets which class counts as positive, and metrics such as precision and recall depend on that choice.

    Class imbalance is handled with weights. sample_weight_col names a column of per-row weights. For classification, every sample in a class must have the same weight, so the weights act as per-class importance. Classes with higher weights have more influence on the learning algorithm.

    AutoML parameters that shape what training optimises and how data is split
    ParameterWhat it controls
    primary_metricMetric used to evaluate and rank models (classification default f1, regression default r2)
    pos_labelPositive class for binary classification; affects precision and recall
    sample_weight_colPer-row weights; for classification, one weight per class, from 0 to 10,000
    time_colChronological split: earliest points train, latest points test
    split_colUser-specified train / validate / test assignment per row

    Checkpoint 3 of 4· Exam question

    A data scientist is building a training pipeline for a scikit-learn classifier on Databricks. To evaluate the model, they fit `StandardScaler` on the entire training DataFrame, transform it, and then run `cross_val_score` on the classifier using the scaled features. A colleague reviewing the pipeline flags this as a data leakage risk. Which change most directly fixes the leakage while keeping the workflow inside a single reusable notebook cell?

    Checkpoint 4 of 4· Match them up

    Match each AutoML parameter to its role

    Tap a term, then the definition that fits it.

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Every step in a scikit-learn Pipeline must have a predict method.Why is that wrong?

      Only the final step may be something other than a transformer, and even that step doesn't need predict. The pipeline simply exposes whatever methods its last step provides.

      Covered in Implementing the training step with a scikit-learn Pipeline

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Evaluate model quality by testing on held-out data.”
      ↩︎ What a training pipeline contains
      “The training process logs model parameters, metrics, and artifacts to the MLflow Tracking server.”
      ↩︎ Key concept
      “The purpose of evaluation is to determine if the newly developed model performs better than the current production model.”
      ↩︎ Checkpoint
    2. 2.
      “A run is a single execution of model code.”
      ↩︎ What a training pipeline contains
      “To control variation caused by differences in ordering and partitioning, use the PySpark functions repartition and sortWithinPartitions.”
      ↩︎ What a training pipeline contains
      “This example notebook shows how to use autologging with scikit-learn.”
      ↩︎ Implementing the training step with a scikit-learn Pipeline
    3. 3.
      “You only have to call fit and predict once on your data to fit a whole sequence of estimators.”
      ↩︎ Implementing the training step with a scikit-learn Pipeline
      “TransformedTargetRegressor deals with transforming the target (i.e. log-transform y).”
      ↩︎ Implementing the training step with a scikit-learn Pipeline
      “A pipeline exposes all methods provided by the last estimator”
      ↩︎ Exam trap 1
      “Pipelines require all steps except the last to be a transformer.”
      ↩︎ Checkpoint
    4. 4.
      “These weights adjust the importance of each class during model training.”
      ↩︎ Choosing a metric and weighting imbalanced classes
      “The positive class. This is useful for calculating metrics such as precision and recall.”
      ↩︎ Choosing a metric and weighting imbalanced classes
      “Supported metrics for classification: “f1” (default), “log_loss”, “precision”, “accuracy”, “roc_auc””
      ↩︎ Prediction
      “Classes with higher sample weights are considered more important, and have a greater influence on the learning algorithm.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Hyperparameter Tuning on Databricks: Cross-Validation, Hyperopt fmin, and SparkTrials

    Spotted a mistake, or was something unclear? Tell us.