What you will be able to do
- Work out how many models a grid search with k-fold cross-validation trains by multiplying the number of hyperparameter combinations by the number of folds
- Turn a parameter grid into its number of combinations by multiplying the number of values listed for each hyperparameter
- Count the model fits when a Hyperopt objective runs cross-validation inside every trial, and tell apart settings that change the total count from settings that only change speed
Key concept
Models trained = combinations × folds — In k-fold cross-validation, every candidate hyperparameter setting is fitted k separate times, once for each fold. A tuning run therefore trains (number of candidate settings) × k models. Scikit-learn's GridSearchCV also refits the winning setting once more by default.
1.Cross-validation trains one new model per fold
To count the models in a tuning job, first be clear about what cross-validation does. In k-fold cross-validation, the training data is split into k parts called folds. One at a time, each fold is held out. A model is trained on the other k−1 folds and then scored on the held-out fold. The score reported at the end is the average of those k scores. So k-fold cross-validation does not train one model and score it k times. It trains k separate models on k different training sets, which is why the scikit-learn guide warns that the approach "can be computationally expensive."
The scikit-learn guide's own example shows this. It calls cross_val_score(clf, X, y, cv=5) on a linear SVM and describes the call as splitting the data, fitting a model and computing the score five times, with a different split each time. The result is an array of five scores, one for each fitted model. If you leave out cv, cross_val_score uses 5 folds by default in current scikit-learn versions.
Checkpoint 1 of 6· Check yourself
You run cross_val_score(clf, X, y, cv=5) on a single estimator with fixed hyperparameters. How many times is the estimator fitted?
Each fold gets its own model, trained on the other four folds. With cv=5, that makes five fits and five scores. cross_val_score does not refit a final model.
“fitting a model and computing the score 5 consecutive times (with different splits each time)”Source: scikit-learn.org
The Databricks Hyperopt documentation uses the same unit. It says that in Hyperopt, "a trial generally corresponds to fitting one model on one setting of hyperparameters." Keep that unit in mind: one model is one hyperparameter setting fitted on one training set. A fold is a different training set, so it means another fit.
2.A grid is a product, so the counts multiply
Grid search tries every combination of the values you list. If you list 3 values for one hyperparameter and 3 for another, the grid does not hold 3 + 3 = 6 candidates. It holds 3 × 3 = 9. Each extra hyperparameter multiplies the total again by the number of values listed for it.
The full formula is: models trained = (product of the number of values for each hyperparameter) × (number of folds). For the grid above, that is 3 × 3 × 5 = 45. In the Pipeline example the parameters belong to different steps (reduce_dim__ for PCA and clf__ for the classifier). That does not change the count, because every combination is still one candidate. The Databricks Hyperopt docs make the same equation between candidate settings and fits: max_evals is the "number of hyperparameter settings to try (the number of models to fit)". With cross-validation added, each of those settings costs k fits instead of one.
3 × 2 × 3 = 18 combinations. clf__C applies to whichever classifier is chosen, so it multiplies both. 18 × 5 folds = 90 fits. GridSearchCV's default refit of the best combination on the full training set adds 1, for 91.
Checkpoint 2 of 6· Exam question
A data scientist configures a scikit-learn `GridSearchCV` on Databricks with a parameter grid containing `n_estimators: [100, 200, 300]` and `max_depth: [4, 8]`, using `cv=5` and the default `refit=True`. How many total model fits does this search perform, including the final refit on the full training set?
Correct answer: A — 31 fits, from 6 hyperparameter combinations evaluated across 5 folds plus one refit on the entire training set
- A. The grid has 3 values for `n_estimators` and 2 for `max_depth`, giving 6 combinations, and each combination is trained and validated once per fold under `cv=5`, producing 6 x 5 = 30 fits during the search. With the default `refit=True`, scikit-learn then trains one additional model on the entire training set using the best-found parameters, bringing the total to 31.
- B. This correctly computes the 30 search fits from 6 combinations times 5 folds, but it wrongly assumes no additional training happens afterward. `GridSearchCV` defaults to `refit=True`, so a final model is always retrained on the full dataset unless the caller explicitly passes `refit=False`.
- C. Adding combinations and folds instead of multiplying them misrepresents how cross-validated grid search works. Every hyperparameter combination is evaluated against every fold independently, so the counts multiply rather than sum, making this total far too low.
- D. This treats grid search as if one model were fit per combination and simply scored differently per fold, but that is not how cross-validation works. Each fold requires training a fresh model on that fold's training split, so the fold count is a true multiplier, not a reused artifact.
3.Cross-validation inside a Hyperopt objective
The same multiplication applies when the search is not a grid. The Databricks example for parallelizing Hyperopt tunes the C parameter of an SVM. Each candidate's score comes from cross-validation, not from a single split. (Databricks notes that Hyperopt is no longer included in Databricks Runtime ML after 16.4 LTS ML. The counting logic is the same for any tuner.)
def objective(C):
# Create a support vector classifier model
clf = SVC(C=C)
# Use the cross-validation accuracy to compare the models' performance
accuracy = cross_val_score(clf, X, y).mean()
# Hyperopt tries to minimize the objective function. A higher accuracy value means a better model, so you must return the negative accuracy.
return {'loss': -accuracy, 'status': STATUS_OK}The tuning loop sets how many candidates are tried, and the objective function sets how many models each candidate costs. The Databricks page tells you to "set max_evals to the maximum number of points in hyperparameter space to test". The example uses max_evals=16. Inside each trial, cross_val_score(clf, X, y) is called without cv, so it uses the default 5 folds. The run therefore tries 16 candidate values of C and trains up to 16 × 5 = 80 models. max_evals counts candidate settings, not individual fits.
Checkpoint 3 of 6· Fill the gap
Which fmin() argument sets how many hyperparameter settings Hyperopt will try?
argmin = fmin(
fn=objective,
space=search_space,
algo=algo,
? =16)max_evals is the maximum number of hyperparameter settings to test. max_queue_len controls how many settings are generated ahead of time, parallelism is an argument of SparkTrials, and n_trials is an Optuna argument.
Source: docs.databricks.comDistributing the run does not change the count. When you pass SparkTrials to fmin(), its parallelism argument is the "number of models to fit and evaluate concurrently". That controls how many trials run at the same time, which changes how long the run takes. It does not change how many models are trained in total.
Checkpoint 4 of 6· Check yourself
A Hyperopt run uses max_evals=16, an objective that runs 5-fold cross_val_score, and SparkTrials(parallelism=4). Which statement is correct?
Each of the up to 16 trials runs 5-fold cross-validation, which gives up to 80 fits. parallelism only sets how many trials run at the same time.
“parallelism: Number of models to fit and evaluate concurrently.”Source: docs.databricks.com
Checkpoint 5 of 6· Exam question
A team is comparing two hyperparameter search strategies for a classification model that takes a long time to train. They must estimate compute cost before committing to a Databricks cluster size. Which pairing of setups results in the larger total number of model fits during the search phase (excluding any refit)?
Correct answer: A — `GridSearchCV` over a grid with 8 total combinations using 10-fold cross-validation, producing 80 search fits
- A. Grid search evaluates every one of the 8 parameter combinations against each of the 10 folds, so the search performs 8 x 10 = 80 fits, which is the largest total among the choices presented here.
- B. Randomized search multiplies its sampled iterations by the number of folds in exactly the same way grid search does, so 8 iterations across 10 folds also produces 8 x 10 = 80 fits, tying rather than exceeding the grid search total in the first option.
- C. Reducing the fold count to 5 while keeping 8 combinations yields 8 x 5 = 40 fits, which is half of the 80 fits produced when the same grid is paired with 10-fold cross-validation.
- D. With 5 sampled iterations across 8 folds, the total is 5 x 8 = 40 fits, matching the reduced-fold grid search option but still well below the 80 fits produced by either setup that combines 8 candidates with 10 folds.
Sources4
4.Counting models on exam questions
Every counting question in this subdomain can be solved in three steps. First, count the candidate settings: multiply the number of values listed for each hyperparameter, or read max_evals for Hyperopt. Second, find k, the number of folds. Use the stated cv value or the tool's default: 5 for scikit-learn's GridSearchCV and cross_val_score, and 3 for Spark MLlib's CrossValidator (numFolds). Third, multiply the two. Add one only if the question asks about the final refit model. A TrainValidationSplit is the single-split case: it trains one model per setting instead of k.
| Search | Combinations | Folds (cv) | Models during search | With GridSearchCV refit |
|---|---|---|---|---|
| param_grid: reduce_dim__n_components (3 values) × clf__C (3 values) | 9 | 5 | 45 | 46 |
| param_grid: reduce_dim (3) × clf (2) × clf__C (3) | 18 | 5 | 90 | 91 |
| param_grid: clf__C (3 values) only | 3 | 5 | 15 | 16 |
Checkpoint 6 of 6· Exam question
A data scientist runs `GridSearchCV` with a grid producing 12 hyperparameter combinations, `cv=4`, and `refit=False` because a separate holdout set will be used to pick the deployment model. How many total model fits occur during this search?
Correct answer: A — 48 fits, since the 12 combinations are each trained once per fold and no refit model is added because `refit` is disabled
- A. Every one of the 12 hyperparameter combinations is trained and validated once for each of the 4 folds, giving 12 x 4 = 48 fits. Setting `refit=False` only removes the final full-dataset retraining step, so the search-phase count itself is unaffected.
- B. This assumes `refit=False` is ignored, but scikit-learn's `GridSearchCV` respects that flag explicitly and skips the final full-dataset model whenever it is set. No hidden extra fit occurs beyond the 48 search fits, so this total overstates the work performed.
- C. This confuses `refit=False` with skipping cross-validation entirely, but the flag only controls the optional final model trained on the full dataset. Every fold still requires training a separate model per combination, so far more than 12 fits occur.
- D. Combining 12 combinations and 4 folds by addition rather than multiplication misrepresents the nested loop structure of cross-validated grid search, where every combination is retrained from scratch on each fold's training split.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The number of models trained equals the number of hyperparameter combinations in the grid.Why is that wrong?
Each combination is cross-validated, so it is fitted once per fold. The total is combinations × folds, plus one refit in GridSearchCV's default setup.
Covered in A grid is a product, so the counts multiply
2.k-fold cross-validation trains one model and scores it on k different subsets.Why is that wrong?
Each fold trains a new model on the other k−1 folds, and that model is scored only on its held-out fold. k folds means k separate fits.
3.In Hyperopt, max_evals is always the total number of models trained.Why is that wrong?
max_evals counts hyperparameter settings (trials). If the objective function runs cross-validation, each trial fits one model per fold, so the total is max_evals × folds.
Covered in Cross-validation inside a Hyperopt objective
4.Raising SparkTrials parallelism changes the number of models trained.Why is that wrong?
parallelism only sets how many trials run at the same time. The total number of fits is still decided by the number of settings tried and the folds per setting.
Covered in Cross-validation inside a Hyperopt objective
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-conceptsOfficial docs
“In Hyperopt, a trial generally corresponds to fitting one model on one setting of hyperparameters.”
↩︎ Cross-validation trains one new model per fold“Number of hyperparameter settings to try (the number of models to fit).”
↩︎ A grid is a product, so the counts multiply“Number of hyperparameter settings to try (the number of models to fit).”
↩︎ Counting models on exam questions - 2.https://scikit-learn.org/stable/modules/cross_validation.htmlSecondary source
“The performance measure reported by k-fold cross-validation is then the average of the values computed in the loop.”
↩︎ Cross-validation trains one new model per fold“This approach can be computationally expensive”
↩︎ Cross-validation trains one new model per fold“fitting a model and computing the score 5 consecutive times (with different splits each time)”
↩︎ Key concept“fitting a model and computing the score 5 consecutive times (with different splits each time)”
↩︎ Exam trap 1“the resulting model is validated on the remaining part of the data”
↩︎ Exam trap 2“fitting a model and computing the score 5 consecutive times (with different splits each time)”
↩︎ Checkpoint“fitting a model and computing the score 5 consecutive times (with different splits each time)”
↩︎ Prediction - 3.https://scikit-learn.org/stable/modules/compose.htmlSecondary source
“You can grid search over parameters of all estimators in the pipeline at once.”
↩︎ A grid is a product, so the counts multiply - 4.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-spark-mlflow-integrationOfficial docs
“Set max_evals to the maximum number of points in hyperparameter space to test, that is, the maximum number of models to fit and evaluate.”
↩︎ Cross-validation inside a Hyperopt objective“parallelism: Number of models to fit and evaluate concurrently.”
↩︎ Cross-validation inside a Hyperopt objective“Use the cross-validation accuracy to compare the models' performance”
↩︎ Exam trap 3“parallelism: Number of models to fit and evaluate concurrently.”
↩︎ Exam trap 4 - 5.
“when you run tuning code that uses CrossValidator or TrainValidationSplit, hyperparameters and evaluation metrics are automatically logged in MLflow.”
↩︎ Counting models on exam questions