CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 36/48

    Cross-Validation in Pipelines, Tuning and Spark

    Perform cross-validation as a part of model fitting.

    10 min read
    2.08% of exam
    6 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Cross-validate a preprocessing-plus-model Pipeline so that preprocessing is learned only from each training fold
    • Use cross-validation as the scoring step inside GridSearchCV and inside a tuning objective function, and count the fits a grid search performs
    • Choose between cross_val_score, cross_validate and cross_val_predict, and pick a splitting strategy that fits the data, including time series and Spark

    1.Keep preprocessing inside the folds

    A model must be scored on data it was not trained on, and the same rule applies to preprocessing. Steps such as standardisation or feature selection must be learned from the training data only and then applied to the held-out data. With one split you can do this by hand: fit a StandardScaler on X_train, transform X_train, fit the classifier, transform X_test with the same scaler, and score.

    Checkpoint 1 of 6· Put it in order

    Put the manual, leakage-free sequence for one train/test split in order.

    1. 1.Transform X_train and fit the classifier on the result
    2. 2.Transform X_test with the already-fitted scaler
    3. 3.Fit the StandardScaler on X_train
    4. 4.Score the classifier on the transformed X_test

    Under k-fold CV that sequence has to run once per fold, and a Pipeline does it for you. Pass the whole pipeline to cross_val_score as the estimator. On each fold the scaler is then refitted on that fold's training rows only. In the example below, cv is the five-split ShuffleSplit(n_splits=5, test_size=0.3, random_state=0) iterator defined earlier in the same guide.

    Cross-validating a scaler + SVM pipeline so that scaling is learned inside each foldpython
    >>> from sklearn.pipeline import make_pipeline
    >>> clf = make_pipeline(preprocessing.StandardScaler(), svm.SVC(C=1))
    >>> cross_val_score(clf, X, y, cv=cv)
    array([0.977, 0.933, 0.955, 0.933, 0.977])

    Spark ML combines the same pieces. On Databricks, the pyspark.ml.connect module provides feature transformers, ML pipelines and cross-validation together.

    Checkpoint 2 of 6· Exam question

    A binary classification dataset is severely imbalanced, with 95% negative and 5% positive labels. An engineer wants to run k-fold cross-validation while keeping the same 95/5 class ratio inside every fold. Which cross-validator should they use?

    Sources12

    2.Cross-validation as the scoring step in tuning

    In practice you rarely cross-validate a single configuration. CV is usually the step that scores each candidate during hyperparameter tuning. GridSearchCV builds this in: you give it an estimator, a param_grid, a scoring rule and cv, and it cross-validates every combination in the grid.

    GridSearchCV: two values of C, each scored by 5-fold CV with a custom F2 scorerpython
    >>> from sklearn.metrics import fbeta_score, make_scorer
    >>> ftwo_scorer = make_scorer(fbeta_score, beta=2)
    >>> from sklearn.model_selection import GridSearchCV
    >>> from sklearn.svm import LinearSVC
    >>> grid = GridSearchCV(LinearSVC(), param_grid={'C': [1, 10]},
    ...                     scoring=ftwo_scorer, cv=5)

    Checkpoint 3 of 6· Check yourself

    Each k-fold run trains one model per fold. How many models does the search above train while cross-validating the grid?

    You can also put CV inside a tuning objective yourself. Databricks' Hyperopt example does this: the objective function scores each candidate C by its mean cross-validation accuracy. Because the optimiser minimises its objective, the function returns the accuracy negated.

    Cross-validation inside a tuning objective: the mean CV accuracy, negated, becomes the losspython
    def objective(C):
        # Create a support vector classifier model
        clf = SVC(C=C)
    
        # Use the cross-validation accuracy to compare the models' performance
        accuracy = cross_val_score(clf, X, y).mean()
    
        # Hyperopt tries to minimize the objective function. A higher accuracy value means a better model, so you must return the negative accuracy.
        return {'loss': -accuracy, 'status': STATUS_OK}

    Checkpoint 4 of 6· Exam question

    A team configures `GridSearchCV` with a `param_grid` covering 3 values for `max_depth` and 4 values for `n_estimators`, and sets `cv=5`. How many individual model fits does this search perform in total?

    Sources34

    3.More than one number: cross_validate and cross_val_predict

    cross_val_score returns one metric per fold. cross_validate differs in two ways: it accepts several metrics at once, and it returns a dict that also contains fit times and score times. Training-set scores are left out by default to save compute. Set return_train_score=True to get them, which lets you compare train and test scores fold by fold. return_estimator=True keeps each fold's fitted model, and return_indices=True keeps each fold's train/test indices.

    cross_validate with two metrics returns one test_<metric> key per scorer, plus timingspython
    >>> from sklearn.model_selection import cross_validate
    >>> from sklearn.metrics import recall_score
    >>> scoring = ['precision_macro', 'recall_macro']
    >>> clf = svm.SVC(kernel='linear', C=1, random_state=0)
    >>> scores = cross_validate(clf, X, y, scoring=scoring)
    >>> sorted(scores.keys())
    ['fit_time', 'score_time', 'test_precision_macro', 'test_recall_macro']

    cross_val_predict takes the same arguments but returns something different. For every input row, it returns the prediction that row received while it was in a test fold. These predictions come from several different models mixed together, so they are not a valid estimate of generalisation error. Use them to visualise predictions or for model blending.

    The three scikit-learn cross-validation helpers compared
    HelperReturnsUse it for
    cross_val_scoreAn array with one score per foldA single-metric generalisation estimate (mean and standard deviation)
    cross_validateA dict with test_<scorer> arrays, fit_time and score_time; optionally train scores, estimators and indicesSeveral metrics at once, timing, and train-versus-test comparisons
    cross_val_predictOne out-of-fold prediction per input rowVisualisation and model blending, not error estimation

    Checkpoint 5 of 6· Check yourself

    A colleague computes accuracy from the output of cross_val_predict and reports it as the model's expected accuracy on new data. What is wrong with that?

    Sources52

    4.Matching the split to the data: time series and Spark

    Shuffled k-fold assumes the rows are independent and identically distributed. That assumption breaks when the data was produced over time. Then a time-series-aware scheme is safer, so that the model is never validated on data older than what it trained on. Databricks AutoML follows this for forecasting. It uses time series cross-validation, which extends the training window forward in time and validates on the time points that come after it. The number of folds depends on the table, for example how many series there are and how long they are.

    For Spark ML on Databricks, pyspark.ml.connect provides a CrossValidator under pyspark.ml.connect.tuning. The older MLlib automated MLflow tracking logged CrossValidator and TrainValidationSplit runs automatically. It is deprecated and disabled by default on Databricks Runtime 10.4 LTS ML and above. Use MLflow PySpark ML autologging, mlflow.pyspark.ml.autolog(), instead. It is enabled by default through Databricks Autologging.

    Checkpoint 6 of 6· Match them up

    Match each situation to the cross-validation approach the sources point to.

    Tap a term, then the definition that fits it.

    Sources6512

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Standardising the full dataset once and then cross-validating only the model is fine, because scaling isn't learning.Why is that wrong?

      Preprocessing learns parameters from data, so it must be fitted on each training fold only. Wrapping it in a Pipeline gives you that behaviour under cross-validation.

      Covered in Keep preprocessing inside the folds

    2. 2.A metric computed on cross_val_predict output is the same as the cross-validated score.Why is that wrong?

      cross_val_score averages per-fold scores. cross_val_predict pools predictions from different models and is meant for visualisation or blending, not for estimating generalisation error.

      Covered in More than one number: cross_validate and cross_val_predict

    3. 3.A tuning objective can return the mean cross-validation accuracy directly as its loss.Why is that wrong?

      The optimiser minimises the objective, so returning accuracy as the loss would select the worst model. Return the accuracy negated.

      Covered in Cross-validation as the scoring step in tuning

    Practise it for real

    Cross-validate an SVM on the iris dataset three ways (plain, pipelined, multi-metric) and compare what each call returns.

    1. 1.Load iris with datasets.load_iris(return_X_y=True), build svm.SVC(kernel='linear', C=1, random_state=42) and run cross_val_score(clf, X, y, cv=5). Print scores.mean() and scores.std().

      Why: This gives you a baseline 5-fold estimate and shows how much it varies from fold to fold.

      You should see: Five per-fold scores close to array([0.96, 1. , 0.96, 0.96, 1. ]), about 0.98 accuracy with a standard deviation of 0.02.

    2. 2.Define cv = ShuffleSplit(n_splits=5, test_size=0.3, random_state=0), wrap the model as make_pipeline(preprocessing.StandardScaler(), svm.SVC(C=1)) and call cross_val_score(clf, X, y, cv=cv).

      Why: The scaler is now refitted inside every split, so no statistics from the held-out rows reach training.

      You should see: Five scores like array([0.977, 0.933, 0.955, 0.933, 0.977]).

    3. 3.Call cross_validate(clf, X, y, scoring=['precision_macro', 'recall_macro']) on svm.SVC(kernel='linear', C=1, random_state=0) and print sorted(scores.keys()).

      Why: This shows that cross_validate returns several metrics and timings in one call.

      You should see: ['fit_time', 'score_time', 'test_precision_macro', 'test_recall_macro']

    Stuck? Get a nudge

    Add return_train_score=True to the last call and compare the train_ and test_ arrays fold by fold.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “including classification, feature transformers, ML pipelines, and cross validation”
      ↩︎ Keep preprocessing inside the folds
      “Model tuning: pyspark.ml.connect.tuning.CrossValidator”
      ↩︎ Matching the split to the data: time series and Spark
    2. 2.
      “A Pipeline makes it easier to compose estimators, providing this behavior under cross-validation”
      ↩︎ Keep preprocessing inside the folds
      “It allows specifying multiple metrics for evaluation.”
      ↩︎ More than one number: cross_validate and cross_val_predict
      “cross_val_predict is not an appropriate measure of generalization error.”
      ↩︎ More than one number: cross_validate and cross_val_predict
      “If one knows that the samples have been generated using a time-dependent process, it is safer to use a time-series aware cross-validation scheme.”
      ↩︎ Matching the split to the data: time series and Spark
      “A Pipeline makes it easier to compose estimators, providing this behavior under cross-validation”
      ↩︎ Exam trap 1
      “cross_val_predict is not an appropriate measure of generalization error.”
      ↩︎ Exam trap 2
      “similar data transformations similarly should be learnt from a training set and applied to held-out data for prediction”
      ↩︎ Checkpoint
      “The performance measure reported by k-fold cross-validation is then the average of the values computed in the loop.”
      ↩︎ Checkpoint
    3. 3.
      “Hyperopt tries to minimize the objective function. A higher accuracy value means a better model, so you must return the negative accuracy.”
      ↩︎ Cross-validation as the scoring step in tuning
      “A higher accuracy value means a better model, so you must return the negative accuracy.”
      ↩︎ Exam trap 3
    4. 4.
      “take a scoring parameter that controls what metric they apply to the estimators evaluated.”
      ↩︎ Cross-validation as the scoring step in tuning
    5. 5.
      “hyperparameters and evaluation metrics are automatically logged in MLflow.”
      ↩︎ More than one number: cross_validate and cross_val_predict
      “MLlib automated MLflow tracking is deprecated and disabled by default on clusters that run Databricks Runtime 10.4 LTS ML and above.”
      ↩︎ Matching the split to the data: time series and Spark
    6. 6.
      “This method incrementally extends the training dataset chronologically and performs validation on subsequent time points.”
      ↩︎ Matching the split to the data: time series and Spark
      “For forecasting tasks, AutoML uses time series cross-validation.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise Databricks Certified Machine Learning Associate in quiz mode.

    Spotted a mistake, or was something unclear? Tell us.