What you will be able to do
- Explain why a single train-validation split gives a score that depends on which rows happened to land in the validation set
- Describe how k-fold cross-validation uses every training row for both fitting and validation, and reports an averaged score
- State the main downside of cross-validation: each candidate model must be fit k times instead of once
- Choose between a single split, k-fold cross-validation and time-series cross-validation for a given dataset size, compute budget and data shape
Key concept
k-fold cross-validation — Instead of fixing one validation set, you split the training data into k folds. You train k times, each time holding out a different fold, and average the k validation scores. You get a lower-variance estimate and no data is wasted, but it costs k times the training compute.
1.The single train-validation split and its hidden weakness
To tune hyperparameters or compare models, you need a score measured on data the model did not train on. The simplest way to get one is to carve the data into fixed pieces. You train on one piece, compare candidates on a validation piece, and keep a test piece for the final, untouched evaluation. Databricks AutoML does exactly this by default: when you don't specify a split strategy, it partitions the table into training, validation and test sets.
That split is cheap. Each candidate model is fit once and scored once. But it has two costs that are easy to miss. First, every row you reserve for validation is a row the model never learns from. With a 60/20/20 split, the model sees only 60% of your labelled data. Second, the validation score comes from one particular random set of rows. Reshuffle with a different seed and a different 20% lands in validation, so the score moves. On a large dataset that movement is small. On a small dataset it can be big enough to change which hyperparameter setting looks best.
Checkpoint 1 of 8· Check yourself
Why can the validation score from a single train-validation split be unreliable for choosing between hyperparameter settings?
A single split fixes one random choice of validation rows, so a different shuffle can give a different score and change which setting looks best.
“the results can depend on a particular random choice for the pair of (train, validation) sets.”Source: scikit-learn.org
Nothing went wrong. That is how a single split behaves. The 400 validation rows are one random sample, and a 0.01 gap is within the noise you get from picking a different sample. The validation score depends on which rows ended up in the validation set, so a single split can't reliably separate two settings that close together.
2.How k-fold cross-validation fixes it
Cross-validation replaces the fixed validation set with rotation. The training data is split into k folds. For each fold, a model is trained on the other k-1 folds and scored on that one. After k rounds, every training row has been used for validation exactly once and for fitting k-1 times. The reported score is the average of the k fold scores, and their spread (standard deviation) tells you how stable the estimate is. That fixes both weaknesses of the single split. No rows are permanently given up to validation, and the result no longer depends on one lucky or unlucky set of rows.
>>> from sklearn.model_selection import cross_val_score
>>> clf = svm.SVC(kernel='linear', C=1, random_state=42)
>>> scores = cross_val_score(clf, X, y, cv=5)
>>> scores
array([0.96, 1. , 0.96, 0.96, 1. ])One thing cross-validation does not replace is the test set. The folds take over the validation role, but you still keep a test set that neither training nor tuning ever touches, so the final number reports generalization rather than how well you tuned. The same idea shows up in Databricks AutoML for forecasting. There the folds respect time order, which gives a robust evaluation across different segments of the series rather than a single cut.
Checkpoint 2 of 8· Check yourself
A team switches from a train/validation/test split to 5-fold cross-validation for tuning. Which set can they now drop?
The rotating folds replace the fixed validation set, but you still keep a test set for the final, unbiased evaluation.
“A test set should still be held out for final evaluation, but the validation set is no longer needed when doing CV.”Source: scikit-learn.org
Checkpoint 3 of 8· Exam question
A data scientist has 800 labeled rows for a churn model, and a single 80/20 train-validation split produces a validation accuracy that swings between 0.71 and 0.83 depending on the random seed used for the split. The team wants a more trustworthy estimate of how the model will generalize before presenting results to stakeholders. Which change addresses this instability?
Correct answer: A — Switch to `cross_val_score` with 5-fold cross-validation so the reported metric averages five different train/test partitions instead of one arbitrary split.
- A. Five-fold cross-validation trains and evaluates the model on five distinct partitions of the data and averages the resulting scores, which directly reduces the variance caused by depending on one arbitrary split and gives a more reliable generalization estimate.
- B. Enlarging the validation set shrinks the training set even further on an already small 800-row dataset, which can hurt model quality and does not address the underlying problem that only one random partition is being evaluated.
- C. Fixing the random seed makes the result reproducible but does not make it more reliable — it locks in whichever lucky or unlucky split was drawn instead of averaging across several partitions.
- D. Repeating training on the identical split and keeping only the best score does not reduce variance at all; it introduces selection bias by cherry-picking the most favorable run rather than averaging independent estimates.
3.The price: k fits per candidate
Cross-validation's main downside is compute. A single split fits each candidate model once. k-fold cross-validation fits it k times, once per fold, so for the same number of candidates you pay roughly k times the training time. In a hyperparameter search that multiplier applies to every setting you try. A search that would fit 50 models with a single split fits 250 with 5-fold CV. For a small scikit-learn model that's trivial. For a large gradient-boosted ensemble or a model trained on a big table, it can be the difference between minutes and hours.
On Databricks, cross-validation usually runs inside the tuning loop. In the Hyperopt example below, the objective function scores each proposed value of C with cross_val_score and averages across folds. So every trial Hyperopt counts is really k model fits. The usual way to absorb the extra cost is parallelism. You distribute trials across Spark workers rather than drop cross-validation, but the total amount of training work still grows with k.
def objective(C):
# Create a support vector classifier model
clf = SVC(C=C)
# Use the cross-validation accuracy to compare the models' performance
accuracy = cross_val_score(clf, X, y).mean()
# Hyperopt tries to minimize the objective function. A higher accuracy value means a better model, so you must return the negative accuracy.
return {'loss': -accuracy, 'status': STATUS_OK}Checkpoint 4 of 8· Fill the gap
Which function completes this objective so that each hyperparameter setting is scored by cross-validation rather than a single split?
# Use the cross-validation accuracy to compare the models' performance
accuracy = ? (clf, X, y).mean()cross_val_score returns one score per fold, and .mean() averages them. train_test_split would only produce one split, while fmin and SparkTrials drive the search itself.
Source: docs.databricks.comCheckpoint 5 of 8· Check yourself
Compared with a single train-validation split, what is the main cost of evaluating every candidate with k-fold cross-validation?
Cross-validation trades compute for a better use of data: each candidate is fit once per fold, but no rows are permanently given up to validation.
“This approach can be computationally expensive, but does not waste too much data (as is the case when fixing an arbitrary validation set)”Source: scikit-learn.org
Checkpoint 6 of 8· Exam question
A team is training a gradient-boosted tree model on a dataset with only 400 rows, which is too small to comfortably carve out a separate validation set without starving the training set. Which property of k-fold cross-validation makes it a better fit than a single train-validation split for this specific situation?
Correct answer: A — Every row is used for training in some folds and for testing in exactly one fold, so the full dataset contributes to both training and evaluation instead of being permanently set aside as a validation set.
- A. This is correct: with k folds, each row serves as training data in k-1 folds and as test data in exactly one fold, so no data is permanently sacrificed to a validation set the way it would be with a single split, which matters most when the total row count is already small.
- B. Cross-validation does not generate synthetic rows or otherwise increase the dataset size; it only reorganizes how the existing 400 rows are partitioned across repeated train/test cycles.
- C. Cross-validation is typically used during model selection and hyperparameter tuning, but a separate held-out test set is still recommended for a final, unbiased evaluation of the chosen model; it does not eliminate that need.
- D. Cross-validation requires more model fits, not fewer, since a model is trained once per fold instead of once for a single split, so it increases rather than reduces total training cost for expensive models like gradient-boosted trees.
4.Choosing between them
Neither approach is always right. The choice depends on two questions: how scarce is the data, and how expensive is each fit? When data is small, cross-validation earns its cost. Holding out a fixed validation set would throw away a large share of a small dataset, and a single split's score would be noisy. When data is large and each fit is expensive, a single train-validation split is often good enough. A big validation set already gives a stable score, and multiplying training time by k buys little extra certainty. Spark MLlib offers both approaches as separate tuning classes, CrossValidator and TrainValidationSplit, so you choose based on that tradeoff.
| Aspect | Train-validation split | k-fold cross-validation |
|---|---|---|
| scikit-learn helper | train_test_split | cross_val_score (KFold or StratifiedKFold folds) |
| Fits per candidate | 1 | k |
| Data used for training | Training portion only. Validation rows are never fit | Every row is fit in k-1 of the k rounds |
| Stability of the score | Depends on one random choice of validation rows | Average of k fold scores, with a spread you can inspect |
| Best suited to | Large data, expensive models, tight compute budgets | Small or medium data, where a reliable comparison matters |
Data shape matters too. Randomly shuffled folds assume that rows are interchangeable. With time-ordered data that's false: a random fold would train on the future and validate on the past. For forecasting, Databricks AutoML instead uses time-series cross-validation. The training window grows forward in time, and each round is validated on the period that comes next. You keep the variance-reducing benefit of multiple folds without leaking future information.
Checkpoint 7 of 8· Check yourself
You are validating a daily sales forecasting model. Which validation approach keeps the benefit of multiple folds without training on the future?
Time-series cross-validation grows the training window forward in time and validates on the period that follows. That respects temporal order while still averaging over several segments.
“This method incrementally extends the training dataset chronologically and performs validation on subsequent time points.”Source: docs.databricks.com
Checkpoint 8 of 8· Exam question
A data scientist is deciding between a single train-validation split and 10-fold cross-validation for tuning a random forest on a dataset with 50 million rows, where a single training run already takes several hours on the cluster. Which downside of cross-validation is most relevant to this decision?
Correct answer: A — Cross-validation multiplies total training time by roughly the number of folds, since a separate model must be fit for each fold instead of the single fit required by a train-validation split.
- A. With k-fold cross-validation, a full model fit happens for every fold, so ten folds mean roughly ten times the compute of a single train-validation split; on an already multi-hour training job at 50 million rows, that multiplier is the dominant cost concern.
- B. Fold-level training runs are independent of each other and can be distributed across a cluster (for example with `SparkTrials` or parallel job execution), so lack of parallelism is not an inherent limitation of cross-validation.
- C. Cross-validation does not produce a less accurate estimate on large datasets; if anything, its averaged estimate is at least as reliable as a single split, so accuracy loss is not the tradeoff at play here.
- D. Cross-validation does not require the full dataset to be held in driver memory any more than a single split does — both approaches operate on the same distributed dataset, so memory footprint is not the distinguishing downside.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Cross-validation replaces the test set, so once you use k-fold CV you no longer need a held-out test set.Why is that wrong?
Cross-validation replaces only the validation set. You still need a test set that tuning never touched for the final evaluation.
Covered in How k-fold cross-validation fixes it
2.Cross-validation gives a better estimate at roughly the same cost as a single split.Why is that wrong?
Each candidate is trained once per fold, so compute grows about k-fold. The better estimate is paid for in training time.
Covered in The price: k fits per candidate
3.A single validation score is a stable measure of model quality, so a small gap between two settings means one is really better.Why is that wrong?
A single split's score depends on which rows were randomly chosen for validation, so small gaps can flip when the split changes.
Covered in The single train-validation split and its hidden weakness
4.Shuffled k-fold cross-validation is always the safest choice, including for forecasting data.Why is that wrong?
For time-ordered data, Databricks AutoML uses time-series cross-validation, which trains on earlier data and validates on later time points.
Covered in Choosing between them
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“the dataset is randomly split into 60% train split, 20% validate split, and 20% test split”
↩︎ The single train-validation split and its hidden weakness - 2.https://scikit-learn.org/stable/modules/cross_validation.htmlSecondary source
“the results can depend on a particular random choice for the pair of (train, validation) sets.”
↩︎ The single train-validation split and its hidden weakness“The performance measure reported by k-fold cross-validation is then the average of the values computed in the loop.”
↩︎ How k-fold cross-validation fixes it“This approach can be computationally expensive, but does not waste too much data (as is the case when fixing an arbitrary validation set)”
↩︎ The price: k fits per candidate“which is a major advantage in problems such as inverse inference where the number of samples is very small.”
↩︎ Choosing between them“In the basic approach, called k-fold CV, the training set is split into k smaller sets”
↩︎ Key concept“A test set should still be held out for final evaluation, but the validation set is no longer needed when doing CV.”
↩︎ Exam trap 1“This approach can be computationally expensive, but does not waste too much data (as is the case when fixing an arbitrary validation set)”
↩︎ Exam trap 2“the results can depend on a particular random choice for the pair of (train, validation) sets.”
↩︎ Exam trap 3“A test set should still be held out for final evaluation, but the validation set is no longer needed when doing CV.”
↩︎ Checkpoint - 3.
“Cross-validation provides a robust evaluation of a model's performance over different segments of time.”
↩︎ How k-fold cross-validation fixes it“This method incrementally extends the training dataset chronologically and performs validation on subsequent time points.”
↩︎ Choosing between them“This method incrementally extends the training dataset chronologically and performs validation on subsequent time points.”
↩︎ Exam trap 4 - 4.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-spark-mlflow-integrationOfficial docs
“Use the cross-validation accuracy to compare the models' performance”
↩︎ The price: k fits per candidate - 5.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-conceptsOfficial docs
“Hyperopt is not included in Databricks Runtime for Machine Learning after 16.4 LTS ML.”
↩︎ The price: k fits per candidate - 6.
“when you run tuning code that uses CrossValidator or TrainValidationSplit”
↩︎ Choosing between them