CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 36/48

    k-Fold Cross-Validation with cross_val_score

    Perform cross-validation as a part of model fitting.

    8 min read
    2.08% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why a single train/validation split is a weak basis for comparing models, and what k-fold cross-validation does instead
    • Run k-fold cross-validation with scikit-learn's cross_val_score and read the per-fold scores it returns
    • Control the splitting strategy and the metric through the cv and scoring parameters

    Key concept

    k-fold cross-validation — The training data is split into k folds. Each fold takes one turn as the validation data while a model is trained on the other k−1 folds, and the k scores are averaged into one performance estimate. Every training row is used for both fitting and validating, and no single random split decides the result.

    1.The problem with one validation split

    If you score a model on the same rows it was trained on, the score tells you nothing useful. A model that just memorised the labels would score perfectly and still fail on new data. scikit-learn calls this overfitting. The usual fix is to hold out a test set, for example with train_test_split, and score the model only on that.

    Tuning brings a second problem. If you keep adjusting a hyperparameter such as an SVM's C until the test score peaks, information about the test set leaks into your choices, and the test score stops measuring generalisation. The textbook answer is a third partition, a validation set: tune against the validation set and touch the test set only once at the end. Databricks AutoML takes this approach: it divides your data into training, validation and test splits.

    Three partitions have two costs. Each one takes rows away from training. And the validation score depends on which rows ended up in the validation set, so a different random split could pick a different winning model. Cross-validation (CV) fixes both problems. You still hold out a test set for the final evaluation, but you no longer need a separate validation set.

    Checkpoint 1 of 4· Check yourself

    A team switches from a fixed train/validation/test split to 5-fold cross-validation for model selection. What happens to the test set?

    Sources12

    2.How k-fold cross-validation runs

    In basic k-fold CV, the training set is split into k smaller sets called folds. The loop then runs k times. Each time, a model is trained on k−1 folds and scored on the fold that was left out, using a metric such as accuracy. The score you report is the average of those k values. Databricks' own tuning example uses the same idea: its objective function compares candidate models by their mean cross-validation accuracy, not by a score from one split.

    Choosing k is a trade-off between compute and data. k-fold CV trains one model per fold, so it can be computationally expensive, but it does not waste data the way a fixed validation set does. That matters most when you have few samples. Leave-one-out (LOO) is the extreme case, with k equal to the number of samples n. It builds n models instead of k, which makes it more expensive than k-fold whenever k < n. Its error estimate also often has high variance, because the n models are almost identical to each other. That is why 5 or 10 folds is the usual default.

    Checkpoint 2 of 4· Check yourself

    Why does the scikit-learn guide prefer 5 or 10 folds over leave-one-out for model selection?

    Sources32

    3.Running it: cross_val_score, cv and scoring

    The simplest way to cross-validate in scikit-learn is the cross_val_score helper. You pass an unfitted estimator, the features, the labels, and cv. It splits the data, fits a model, and scores it once per fold, then returns an array with one score per fold.

    5-fold cross-validation of a linear SVM on the iris data: one score per foldpython
    >>> from sklearn.model_selection import cross_val_score
    >>> clf = svm.SVC(kernel='linear', C=1, random_state=42)
    >>> scores = cross_val_score(clf, X, y, cv=5)
    >>> scores
    array([0.96, 1. , 0.96, 0.96, 1. ])

    Report scores.mean() together with scores.std(). In the guide's example that gives 0.98 accuracy with a standard deviation of 0.02. The spread tells you how much the estimate moves from fold to fold. The Databricks tuning example uses the same call, accuracy = cross_val_score(clf, X, y).mean(), to reduce each candidate model to a single number.

    Which splitter runs. When cv is an integer, scikit-learn picks the splitter for you. Regressors get KFold. Classifiers get StratifiedKFold. To use a different strategy, pass a cross-validation iterator instead of an integer:

    Passing a ShuffleSplit iterator as cv: five random 70/30 splits instead of k foldspython
    >>> from sklearn.model_selection import ShuffleSplit
    >>> n_samples = X.shape[0]
    >>> cv = ShuffleSplit(n_splits=5, test_size=0.3, random_state=0)
    >>> cross_val_score(clf, X, y, cv=cv)
    array([0.977, 0.977, 1., 0.955, 1.])

    Which metric is used. By default, each fold is scored with the estimator's own score method. To use a different metric, pass a scorer name to scoring. Scorers follow one rule: higher is always better. For that reason, error metrics are exposed in negated form, for example 'neg_mean_squared_error'. A cross-validated MSE therefore comes back as negative numbers.

    Checkpoint 3 of 4· Fill the gap

    Which parameter makes this call score each fold by macro-averaged F1 instead of the estimator's default score?

    >>> scores = cross_val_score(
    ...     clf, X, y, cv=5,  ? ='f1_macro')

    Checkpoint 4 of 4· Exam question

    A data scientist has only 500 labeled rows and wants a robust, low-variance estimate of a scikit-learn classifier's generalization performance without permanently setting aside data as a fixed holdout. Which approach best fits this goal?

    Sources324

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Once you cross-validate, you no longer need a held-out test set.Why is that wrong?

      Cross-validation replaces the separate validation set used for comparing candidates. A test set is still held out and used once for the final evaluation.

      Covered in The problem with one validation split

    2. 2.An integer cv always means plain KFold, so a classifier's folds can end up with very different class mixes.Why is that wrong?

      With an integer cv, cross_val_score uses StratifiedKFold for classifiers and KFold otherwise.

      Covered in Running it: cross_val_score, cv and scoring

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “AutoML splits your data into three splits for training, validation, and testing.”
      ↩︎ The problem with one validation split
    2. 2.
      “the results can depend on a particular random choice for the pair of (train, validation) sets.”
      ↩︎ The problem with one validation split
      “This situation is called overfitting.”
      ↩︎ The problem with one validation split
      “This approach can be computationally expensive, but does not waste too much data”
      ↩︎ How k-fold cross-validation runs
      “In terms of accuracy, LOO often results in high variance as an estimator for the test error.”
      ↩︎ How k-fold cross-validation runs
      “By default, the score computed at each CV iteration is the score method of the estimator.”
      ↩︎ Running it: cross_val_score, cv and scoring
      “The performance measure reported by k-fold cross-validation is then the average of the values computed in the loop.”
      ↩︎ Key concept
      “A test set should still be held out for final evaluation, but the validation set is no longer needed when doing CV.”
      ↩︎ Exam trap 1
      “cross_val_score uses the KFold or StratifiedKFold strategies by default, the latter being used if the estimator derives from ClassifierMixin”
      ↩︎ Exam trap 2
      “A test set should still be held out for final evaluation, but the validation set is no longer needed when doing CV.”
      ↩︎ Checkpoint
      “As a general rule, most authors and empirical evidence suggest that 5 or 10-fold cross validation should be preferred to LOO.”
      ↩︎ Prediction
    3. 3.
      “Use the cross-validation accuracy to compare the models' performance”
      ↩︎ How k-fold cross-validation runs
      “accuracy = cross_val_score(clf, X, y).mean()”
      ↩︎ Running it: cross_val_score, cv and scoring
    4. 4.
      “All scorer objects follow the convention that higher return values are better than lower return values.”
      ↩︎ Running it: cross_val_score, cv and scoring

    Continue to page 2 of 2

    Cross-Validation in Pipelines, Tuning and Spark

    Spotted a mistake, or was something unclear? Tell us.