What you will be able to do
- Cross-validate a preprocessing-plus-model Pipeline so that preprocessing is learned only from each training fold
- Use cross-validation as the scoring step inside GridSearchCV and inside a tuning objective function, and count the fits a grid search performs
- Choose between cross_val_score, cross_validate and cross_val_predict, and pick a splitting strategy that fits the data, including time series and Spark
1.Keep preprocessing inside the folds
A model must be scored on data it was not trained on, and the same rule applies to preprocessing. Steps such as standardisation or feature selection must be learned from the training data only and then applied to the held-out data. With one split you can do this by hand: fit a StandardScaler on X_train, transform X_train, fit the classifier, transform X_test with the same scaler, and score.
Checkpoint 1 of 6· Put it in order
Put the manual, leakage-free sequence for one train/test split in order.
- 1.Transform X_train and fit the classifier on the result
- 2.Transform X_test with the already-fitted scaler
- 3.Fit the StandardScaler on X_train
- 4.Score the classifier on the transformed X_test
The scaler learns its parameters from the training rows only. The held-out rows are transformed with those parameters and then scored.
“similar data transformations similarly should be learnt from a training set and applied to held-out data for prediction”Source: scikit-learn.org
Under k-fold CV that sequence has to run once per fold, and a Pipeline does it for you. Pass the whole pipeline to cross_val_score as the estimator. On each fold the scaler is then refitted on that fold's training rows only. In the example below, cv is the five-split ShuffleSplit(n_splits=5, test_size=0.3, random_state=0) iterator defined earlier in the same guide.
>>> from sklearn.pipeline import make_pipeline
>>> clf = make_pipeline(preprocessing.StandardScaler(), svm.SVC(C=1))
>>> cross_val_score(clf, X, y, cv=cv)
array([0.977, 0.933, 0.955, 0.933, 0.977])Spark ML combines the same pieces. On Databricks, the pyspark.ml.connect module provides feature transformers, ML pipelines and cross-validation together.
Checkpoint 2 of 6· Exam question
A binary classification dataset is severely imbalanced, with 95% negative and 5% positive labels. An engineer wants to run k-fold cross-validation while keeping the same 95/5 class ratio inside every fold. Which cross-validator should they use?
Correct answer: A — StratifiedKFold
- A. `StratifiedKFold` builds each fold so the proportion of each class matches the proportion in the full dataset, which keeps the rare positive class represented in every fold's train and test split.
- B. Plain `KFold` splits rows into folds without regard to class labels, so with a 95/5 imbalance it can easily produce folds where the minority class is underrepresented or missing entirely.
- C. `ShuffleSplit` generates random train/test partitions with possible overlap between iterations and does not account for class proportions, so it offers no built-in guarantee of preserving the 95/5 ratio.
- D. `LeaveOneOut` holds out a single row per iteration and trains on the rest, which is computationally expensive on larger data and does nothing to balance class proportions within folds.
2.Cross-validation as the scoring step in tuning
In practice you rarely cross-validate a single configuration. CV is usually the step that scores each candidate during hyperparameter tuning. GridSearchCV builds this in: you give it an estimator, a param_grid, a scoring rule and cv, and it cross-validates every combination in the grid.
>>> from sklearn.metrics import fbeta_score, make_scorer
>>> ftwo_scorer = make_scorer(fbeta_score, beta=2)
>>> from sklearn.model_selection import GridSearchCV
>>> from sklearn.svm import LinearSVC
>>> grid = GridSearchCV(LinearSVC(), param_grid={'C': [1, 10]},
... scoring=ftwo_scorer, cv=5)Checkpoint 3 of 6· Check yourself
Each k-fold run trains one model per fold. How many models does the search above train while cross-validating the grid?
Each of the 2 candidate values of C goes through 5-fold CV, and each fold trains one model, so the search trains 2 × 5 = 10 models.
“The performance measure reported by k-fold cross-validation is then the average of the values computed in the loop.”Source: scikit-learn.org
You can also put CV inside a tuning objective yourself. Databricks' Hyperopt example does this: the objective function scores each candidate C by its mean cross-validation accuracy. Because the optimiser minimises its objective, the function returns the accuracy negated.
def objective(C):
# Create a support vector classifier model
clf = SVC(C=C)
# Use the cross-validation accuracy to compare the models' performance
accuracy = cross_val_score(clf, X, y).mean()
# Hyperopt tries to minimize the objective function. A higher accuracy value means a better model, so you must return the negative accuracy.
return {'loss': -accuracy, 'status': STATUS_OK}Checkpoint 4 of 6· Exam question
A team configures `GridSearchCV` with a `param_grid` covering 3 values for `max_depth` and 4 values for `n_estimators`, and sets `cv=5`. How many individual model fits does this search perform in total?
Correct answer: A — 60 model fits
- A. The grid has 3 x 4 = 12 unique hyperparameter combinations, and each combination is evaluated with a full 5-fold cross-validation, so 12 x 5 = 60 separate model fits are performed before the best combination is selected.
- B. 12 is only the count of unique hyperparameter combinations (3 `max_depth` values times 4 `n_estimators` values); it ignores that each of those combinations is refit once per cross-validation fold.
- C. 15 would follow from multiplying 3 combinations by 5 folds, but it drops the `n_estimators` dimension entirely and undercounts the actual combination grid `GridSearchCV` iterates over.
- D. 20 would follow from multiplying 4 combinations by 5 folds, but it drops the `max_depth` dimension entirely and undercounts the full 3 x 4 combination grid that `GridSearchCV` actually searches.
3.More than one number: cross_validate and cross_val_predict
cross_val_score returns one metric per fold. cross_validate differs in two ways: it accepts several metrics at once, and it returns a dict that also contains fit times and score times. Training-set scores are left out by default to save compute. Set return_train_score=True to get them, which lets you compare train and test scores fold by fold. return_estimator=True keeps each fold's fitted model, and return_indices=True keeps each fold's train/test indices.
>>> from sklearn.model_selection import cross_validate
>>> from sklearn.metrics import recall_score
>>> scoring = ['precision_macro', 'recall_macro']
>>> clf = svm.SVC(kernel='linear', C=1, random_state=0)
>>> scores = cross_validate(clf, X, y, scoring=scoring)
>>> sorted(scores.keys())
['fit_time', 'score_time', 'test_precision_macro', 'test_recall_macro']cross_val_predict takes the same arguments but returns something different. For every input row, it returns the prediction that row received while it was in a test fold. These predictions come from several different models mixed together, so they are not a valid estimate of generalisation error. Use them to visualise predictions or for model blending.
| Helper | Returns | Use it for |
|---|---|---|
| cross_val_score | An array with one score per fold | A single-metric generalisation estimate (mean and standard deviation) |
| cross_validate | A dict with test_<scorer> arrays, fit_time and score_time; optionally train scores, estimators and indices | Several metrics at once, timing, and train-versus-test comparisons |
| cross_val_predict | One out-of-fold prediction per input row | Visualisation and model blending, not error estimation |
Checkpoint 5 of 6· Check yourself
A colleague computes accuracy from the output of cross_val_predict and reports it as the model's expected accuracy on new data. What is wrong with that?
cross_val_score averages per-fold scores. cross_val_predict pools predictions from different models without distinguishing them, so a metric computed from it is not a generalisation estimate.
“cross_val_predict is not an appropriate measure of generalization error.”Source: scikit-learn.org
4.Matching the split to the data: time series and Spark
Shuffled k-fold assumes the rows are independent and identically distributed. That assumption breaks when the data was produced over time. Then a time-series-aware scheme is safer, so that the model is never validated on data older than what it trained on. Databricks AutoML follows this for forecasting. It uses time series cross-validation, which extends the training window forward in time and validates on the time points that come after it. The number of folds depends on the table, for example how many series there are and how long they are.
Shuffling puts future rows into the training folds and past rows into the validation fold, so the model is scored on predicting the past from the future. Time series CV keeps every validation point after the training window, which matches how the model will actually be used.
For Spark ML on Databricks, pyspark.ml.connect provides a CrossValidator under pyspark.ml.connect.tuning. The older MLlib automated MLflow tracking logged CrossValidator and TrainValidationSplit runs automatically. It is deprecated and disabled by default on Databricks Runtime 10.4 LTS ML and above. Use MLflow PySpark ML autologging, mlflow.pyspark.ml.autolog(), instead. It is enabled by default through Databricks Autologging.
Checkpoint 6 of 6· Match them up
Match each situation to the cross-validation approach the sources point to.
Tap a term, then the definition that fits it.
Pick the splitter that fits the data's structure: chronological splits for time series, stratified folds for classifiers, 5–10 folds for i.i.d. data, and Spark's CrossValidator for Spark ML.
“For forecasting tasks, AutoML uses time series cross-validation.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Standardising the full dataset once and then cross-validating only the model is fine, because scaling isn't learning.Why is that wrong?
Preprocessing learns parameters from data, so it must be fitted on each training fold only. Wrapping it in a Pipeline gives you that behaviour under cross-validation.
Covered in Keep preprocessing inside the folds
2.A metric computed on cross_val_predict output is the same as the cross-validated score.Why is that wrong?
cross_val_score averages per-fold scores. cross_val_predict pools predictions from different models and is meant for visualisation or blending, not for estimating generalisation error.
Covered in More than one number: cross_validate and cross_val_predict
3.A tuning objective can return the mean cross-validation accuracy directly as its loss.Why is that wrong?
The optimiser minimises the objective, so returning accuracy as the loss would select the worst model. Return the accuracy negated.
Practise it for real
Cross-validate an SVM on the iris dataset three ways (plain, pipelined, multi-metric) and compare what each call returns.
1.Load iris with datasets.load_iris(return_X_y=True), build svm.SVC(kernel='linear', C=1, random_state=42) and run cross_val_score(clf, X, y, cv=5). Print scores.mean() and scores.std().
Why: This gives you a baseline 5-fold estimate and shows how much it varies from fold to fold.
You should see: Five per-fold scores close to array([0.96, 1. , 0.96, 0.96, 1. ]), about 0.98 accuracy with a standard deviation of 0.02.
2.Define cv = ShuffleSplit(n_splits=5, test_size=0.3, random_state=0), wrap the model as make_pipeline(preprocessing.StandardScaler(), svm.SVC(C=1)) and call cross_val_score(clf, X, y, cv=cv).
Why: The scaler is now refitted inside every split, so no statistics from the held-out rows reach training.
You should see: Five scores like array([0.977, 0.933, 0.955, 0.933, 0.977]).
3.Call cross_validate(clf, X, y, scoring=['precision_macro', 'recall_macro']) on svm.SVC(kernel='linear', C=1, random_state=0) and print sorted(scores.keys()).
Why: This shows that cross_validate returns several metrics and timings in one call.
You should see: ['fit_time', 'score_time', 'test_precision_macro', 'test_recall_macro']
Stuck? Get a nudge
Add return_train_score=True to the last call and compare the train_ and test_ arrays fold by fold.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/train-model/distributed-training/distributed-ml-for-spark-connectOfficial docs
“including classification, feature transformers, ML pipelines, and cross validation”
↩︎ Keep preprocessing inside the folds“Model tuning: pyspark.ml.connect.tuning.CrossValidator”
↩︎ Matching the split to the data: time series and Spark - 2.https://scikit-learn.org/stable/modules/cross_validation.htmlSecondary source
“A Pipeline makes it easier to compose estimators, providing this behavior under cross-validation”
↩︎ Keep preprocessing inside the folds“It allows specifying multiple metrics for evaluation.”
↩︎ More than one number: cross_validate and cross_val_predict“cross_val_predict is not an appropriate measure of generalization error.”
↩︎ More than one number: cross_validate and cross_val_predict“If one knows that the samples have been generated using a time-dependent process, it is safer to use a time-series aware cross-validation scheme.”
↩︎ Matching the split to the data: time series and Spark“A Pipeline makes it easier to compose estimators, providing this behavior under cross-validation”
↩︎ Exam trap 1“cross_val_predict is not an appropriate measure of generalization error.”
↩︎ Exam trap 2“similar data transformations similarly should be learnt from a training set and applied to held-out data for prediction”
↩︎ Checkpoint“The performance measure reported by k-fold cross-validation is then the average of the values computed in the loop.”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-spark-mlflow-integrationOfficial docs
“Hyperopt tries to minimize the objective function. A higher accuracy value means a better model, so you must return the negative accuracy.”
↩︎ Cross-validation as the scoring step in tuning“A higher accuracy value means a better model, so you must return the negative accuracy.”
↩︎ Exam trap 3 - 4.https://scikit-learn.org/stable/modules/model_evaluation.htmlSecondary source
“take a scoring parameter that controls what metric they apply to the estimators evaluated.”
↩︎ Cross-validation as the scoring step in tuning - 5.
“hyperparameters and evaluation metrics are automatically logged in MLflow.”
↩︎ More than one number: cross_validate and cross_val_predict“MLlib automated MLflow tracking is deprecated and disabled by default on clusters that run Databricks Runtime 10.4 LTS ML and above.”
↩︎ Matching the split to the data: time series and Spark - 6.
“This method incrementally extends the training dataset chronologically and performs validation on subsequent time points.”
↩︎ Matching the split to the data: time series and Spark“For forecasting tasks, AutoML uses time series cross-validation.”
↩︎ Checkpoint