What you will be able to do
- Compare k-fold cross-validation with a single validation split and count the models a tuning run fits
- Write a Hyperopt objective, search space and fmin call, and choose between tpe.suggest and rand.suggest
- Decide when to use SparkTrials versus Trials and how parallelism affects results
1.Cross-validation versus a single validation split
Tuning means comparing candidate settings, and each comparison needs a reliable score. k-fold cross-validation splits the data into k folds. Each fold takes one turn as the validation set while the model trains on the others, and the reported score is the average across those k rounds. A fixed validation split fits only one model per setting but holds back data. Cross-validation costs k fits per setting but uses all the data for validation. That gives the counting rule: models fitted = hyperparameter combinations × folds. On Spark ML, the equivalent tuning classes are CrossValidator and TrainValidationSplit. Their automatic MLflow tracking is deprecated in favour of mlflow.pyspark.ml.autolog().
The simplest scikit-learn entry point is cross_val_score:
>>> from sklearn.model_selection import cross_val_score
>>> clf = svm.SVC(kernel='linear', C=1, random_state=42)
>>> scores = cross_val_score(clf, X, y, cv=5)Databricks' own Hyperopt example scores each candidate with cross_val_score inside the objective function, so every trial already averages over folds.
Checkpoint 1 of 6· Exam question
A training pipeline needs to preprocess a DataFrame containing both numeric columns (impute and scale) and categorical columns (one-hot encode) before fitting a `RandomForestClassifier`. Which approach correctly composes this into a single scikit-learn object that can be passed to `GridSearchCV`?
Correct answer: A — Define a `ColumnTransformer` applying an imputer+scaler pipeline to numeric columns and `OneHotEncoder` to categorical columns, as the first step of an outer `Pipeline` ending in the classifier.
- A. A `ColumnTransformer` is built specifically to route different column subsets to different transformers and concatenate the results, and nesting it as the first step of an outer `Pipeline` gives `GridSearchCV` one composite object to fit and score end to end.
- B. One-hot encoding first turns categorical values into binary indicator columns; running an imputer and scaler over those columns afterward misapplies numeric preprocessing to columns whose values are not continuous, which distorts what the encoding represents.
- C. Fitting two pipelines that each only see part of the columns and then combining `predict` outputs does not produce a single feature matrix for the classifier; it discards the classifier's ability to learn from both column types jointly and cannot be scored as one estimator.
- D. Tree-based models like `RandomForestClassifier` can consume one-hot encoded categorical features without issue, so dropping those columns throws away potentially predictive information rather than solving a real compatibility problem.
Checkpoint 2 of 6· Check yourself
Why might you accept the extra compute cost of k-fold cross-validation over a single fixed validation split?
Cross-validation is more expensive because it fits k models per setting, but every row is used for validation in one of the rounds.
“This approach can be computationally expensive, but does not waste too much data (as is the case when fixing an arbitrary validation set)”Source: scikit-learn.org
2.Tuning with Hyperopt fmin
A Hyperopt workflow has four steps. You define a function to minimize, define a search space over hyperparameters, select a search algorithm, and run fmin(). Most of the code lives in the objective function, which usually trains a model and calculates its loss. Hyperopt always minimizes. If your score is "higher is better", such as accuracy, return its negative as the loss.
The search space can mix categorical options with probability distributions such as uniform and log-based ones. The Databricks example tunes an SVM's regularization parameter C:
search_space = hp.lognormal('C', 0, 1.0)You have two main algorithms to choose from. hyperopt.rand.suggest is random search, which samples the space without adapting. hyperopt.tpe.suggest is the Tree of Parzen Estimators, a Bayesian method that picks new settings based on past results. Databricks notes that Bayesian approaches can be much more efficient than grid search and random search. That efficiency lets you explore more hyperparameters and wider ranges. max_evals caps how many settings are tried, which is also the number of models fitted. One quirk: with hp.choice(), Hyperopt returns and logs the *index* of the choice, so use hyperopt.space_eval() to recover the actual value.
Checkpoint 3 of 6· Put it in order
Put the steps of a Hyperopt workflow in order
- 1.Select a search algorithm
- 2.Define a function to minimize
- 3.Define a search space over hyperparameters
- 4.Run the tuning algorithm with Hyperopt fmin()
fmin() needs the objective, the space and the algorithm as arguments, so those three are defined first.
“Run the tuning algorithm with Hyperopt fmin()”Source: docs.databricks.com
Checkpoint 4 of 6· Fill the gap
Which function scores each candidate C in the Databricks objective function?
def objective(C):
# Create a support vector classifier model
clf = SVC(C=C)
# Use the cross-validation accuracy to compare the models' performance
accuracy = ? (clf, X, y).mean()The objective averages cross-validated accuracy for one setting of C. Hyperopt then calls this function once per trial.
Source: docs.databricks.com3.Parallelizing single-node tuning with SparkTrials
SparkTrials distributes a Hyperopt run without any other change to your code. The driver generates trials, and each trial, meaning one model fitted on one setting, runs as a single-task Spark job on a worker. To use it, add one argument to fmin(). Wrapping the call in mlflow.start_run() turns on the automated MLflow tracking:
spark_trials = SparkTrials()
with mlflow.start_run():
argmin = fmin(
fn=objective,
space=search_space,
algo=algo,
max_evals=16,
trials=spark_trials)| Aspect | SparkTrials | Trials |
|---|---|---|
| Use with | Single-machine algorithms such as scikit-learn | Distributed algorithms such as MLlib or Horovod |
| Where trials are evaluated | On Spark worker nodes | On the cluster driver, so the algorithm can distribute itself |
| MLflow logging | Automated tracking when fmin() runs inside mlflow.start_run() | No automatic logging; call MLflow manually |
SparkTrials takes two optional arguments. parallelism sets how many trials run concurrently; it defaults to the number of available executors and is capped at 128. timeout sets the maximum number of seconds fmin() may run. Parallelism has a cost. TPE proposes new trials from past results, so for a fixed max_evals, higher parallelism finishes sooner but gives each proposal fewer completed results to learn from. Don't use SparkTrials on autoscaling clusters: parallelism is fixed when execution starts, so nodes added later go unused. And for very short trials, Spark overhead can cancel out any speedup.
Checkpoint 5 of 6· Exam question
A team wants `GridSearchCV` to search over both `n_components` for a `PCA` step and `C` for an `SVC` step inside the same `Pipeline`, where those steps are named `reduce_dim` and `clf` respectively. Which `param_grid` correctly targets these nested parameters?
Correct answer: A — Use `{'reduce_dim__n_components': [...], 'clf__C': [...]}`, since Pipeline exposes nested step parameters with the `<step_name>__<param>` double-underscore syntax.
- A. Pipeline's `set_params` and `get_params` expose every nested estimator's parameters through the `<step_name>__<param>` convention, and `GridSearchCV` relies on exactly that convention to route grid values to the right step.
- B. The prefix must match the name assigned to the step when the pipeline was constructed, not the underlying estimator's class name; `PCA__n_components` will raise an error unless a step happens to also be named `PCA`.
- C. Bare parameter names are ambiguous once a pipeline has multiple steps, and `GridSearchCV` does not infer which step a bare name belongs to, so this raises an invalid-parameter error at fit time.
- D. A single dot is not the separator scikit-learn's parameter routing recognizes; the library specifically reserves the double underscore for nested parameter access, so a dot-separated key is treated as an unknown parameter.
Checkpoint 6 of 6· Check yourself
With max_evals fixed at 64, a teammate raises SparkTrials parallelism from 8 to 64 to "get the best result". What is the likely trade-off?
Parallelism speeds up the run, but an adaptive algorithm proposes better settings when more past results are available.
“lower parallelism may lead to better results since each iteration has access to more past results.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.SparkTrials is the right way to speed up Hyperopt for any model, including MLlib.Why is that wrong?
SparkTrials is for single-machine libraries such as scikit-learn. With distributed algorithms like MLlib or Horovod, use the default Trials class.
Covered in Parallelizing single-node tuning with SparkTrials
2.The objective function should return accuracy so Hyperopt can find the highest value.Why is that wrong?
fmin minimizes the loss, so a higher-is-better score must be returned as a negative value.
Covered in Tuning with Hyperopt fmin
3.Setting parallelism as high as possible always gives the best tuning result.Why is that wrong?
For a fixed max_evals, higher parallelism is faster, but lower parallelism lets TPE learn from more completed trials.
Covered in Parallelizing single-node tuning with SparkTrials
Practise it for real
Tune an SVM's C parameter with Hyperopt on a Databricks Runtime ML cluster at 16.4 LTS ML or below, and track every trial in MLflow.
1.Write an objective(C) that builds SVC(C=C), scores it with cross_val_score(clf, X, y).mean(), and returns {'loss': -accuracy, 'status': STATUS_OK}.
Why: Hyperopt minimizes loss, so negating accuracy makes "better" mean "lower".
You should see: Calling objective(1.0) returns a dictionary with a negative loss.
2.Define search_space = hp.lognormal('C', 0, 1.0) and set algo=tpe.suggest.
Why: TPE adapts to past results, unlike rand.suggest.
You should see: No output; the space and algorithm are ready to pass to fmin.
3.Run fmin(fn=objective, space=search_space, algo=algo, max_evals=16) without a trials argument.
Why: This gives a single-machine baseline before you distribute anything.
You should see: A printed dictionary with the best value found for C.
4.Create spark_trials = SparkTrials() and rerun fmin inside with mlflow.start_run(): passing trials=spark_trials.
Why: SparkTrials sends trials to the workers, and the run context turns on automated MLflow tracking.
You should see: The notebook experiment shows the runs. For an objective this small, Spark overhead may make it slower than the baseline.
Stuck? Get a nudge
If the distributed run is not faster, look at trial duration: Spark job overhead dominates when each trial lasts only a few seconds.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-spark-mlflow-integrationOfficial docs
“Use the cross-validation accuracy to compare the models' performance”
↩︎ Cross-validation versus a single validation split“hyperopt.rand.suggest: Random search, a non-adaptive approach that samples over the search space”
↩︎ Tuning with Hyperopt fmin“Hyperopt tries to minimize the objective function.”
↩︎ Exam trap 2“Run the tuning algorithm with Hyperopt fmin()”
↩︎ Checkpoint - 2.
“With MLlib automated MLflow tracking, when you run tuning code that uses CrossValidator or TrainValidationSplit, hyperparameters and evaluation metrics are automatically logged in MLflow.”
↩︎ Cross-validation versus a single validation split“Hyperopt is not included in Databricks Runtime for Machine Learning after 16.4 LTS ML.”
↩︎ Tuning with Hyperopt fmin - 3.https://scikit-learn.org/stable/modules/cross_validation.htmlSecondary source
“The performance measure reported by k-fold cross-validation is then the average of the values computed in the loop.”
↩︎ Cross-validation versus a single validation split“by splitting the data, fitting a model and computing the score 5 consecutive times (with different splits each time)”
↩︎ Prediction“This approach can be computationally expensive, but does not waste too much data (as is the case when fixing an arbitrary validation set)”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-conceptsOfficial docs
“Number of hyperparameter settings to try (the number of models to fit).”
↩︎ Tuning with Hyperopt fmin“SparkTrials accelerates single-machine tuning by distributing trials to Spark workers.”
↩︎ Parallelizing single-node tuning with SparkTrials“Use Trials when you call distributed training algorithms such as MLlib methods or Horovod in the objective function.”
↩︎ Exam trap 1“lower parallelism may lead to better results since each iteration has access to more past results.”
↩︎ Exam trap 3“For models created with distributed ML algorithms such as MLlib or Horovod, do not use SparkTrials.”
↩︎ Prediction“lower parallelism may lead to better results since each iteration has access to more past results.”
↩︎ Checkpoint - 5.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-best-practicesOfficial docs
“Bayesian approaches can be much more efficient than grid search and random search.”
↩︎ Tuning with Hyperopt fmin“When you use hp.choice(), Hyperopt returns the index of the choice list.”
↩︎ Tuning with Hyperopt fmin“Do not use SparkTrials on autoscaling clusters.”
↩︎ Parallelizing single-node tuning with SparkTrials - 6.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-distributed-mlOfficial docs
“Databricks does not support automatic logging to MLflow with the Trials class.”
↩︎ Parallelizing single-node tuning with SparkTrials