What you will be able to do
- List the four steps of a Hyperopt workflow in order
- Write an objective function that returns a loss to minimize
- Configure fmin() with a search space, a random or TPE algorithm and max_evals
- Choose between Trials and SparkTrials, and know where Optuna replaces Hyperopt
1.The four-step Hyperopt workflow
Every Hyperopt search has the same four steps. Databricks' example notebook tunes the regularization parameter C of a scikit-learn support vector classifier on the Iris dataset, and its workflow runs in this order: define a function to minimize, define a search space over hyperparameters, select a search algorithm, then run the tuning algorithm with fmin().
Checkpoint 1 of 7· Put it in order
Put the steps of a Hyperopt workflow in order
- 1.Run the tuning algorithm with Hyperopt fmin()
- 2.Define a function to minimize
- 3.Define a search space over hyperparameters
- 4.Select a search algorithm
fmin() comes last because it needs the objective, the space and the algorithm as its arguments.
“Run the tuning algorithm with Hyperopt fmin().”Source: docs.databricks.com
2.The objective function returns a loss
Most of a Hyperopt workflow lives in the objective function. Hyperopt calls it once per trial with a value drawn from the space. The function trains a model, scores it, and returns a loss, either as a scalar or in a dictionary. In the example below, cross-validated accuracy is the score.
def objective(C):
# Create a support vector classifier model
clf = SVC(C=C)
# Use the cross-validation accuracy to compare the models' performance
accuracy = cross_val_score(clf, X, y).mean()
# Hyperopt tries to minimize the objective function. A higher accuracy value means a better model, so you must return the negative accuracy.
return {'loss': -accuracy, 'status': STATUS_OK}Checkpoint 2 of 7· Check yourself
Your objective computes ROC AUC, where higher is better. What should it return to Hyperopt?
Hyperopt always minimizes, so you negate a higher-is-better metric. The Optuna example on Databricks does the same thing: it returns -roc_auc.
“Hyperopt tries to minimize the objective function.”Source: docs.databricks.com
Sources1
3.Search space, algorithm and the fmin() call
The notebook's search space is a single distribution, search_space = hp.lognormal('C', 0, 1.0), and it sets algo=tpe.suggest for Bayesian search. To run random search over the same space, you would pass rand.suggest instead. Nothing else changes, and you can test the same space with either algorithm. The fmin() call then puts the pieces together:
| Argument | What it controls |
|---|---|
| fn | The objective function; returns the loss |
| space | The hyperparameter space: categorical options or distributions such as uniform and log |
| algo | Search algorithm: hyperopt.rand.suggest (random) or hyperopt.tpe.suggest (TPE) |
| max_evals | Number of hyperparameter settings to try, i.e. models to fit |
| trials | A Trials or SparkTrials object |
| early_stop_fn | Optional function deciding whether to stop before max_evals |
Checkpoint 3 of 7· Fill the gap
Which argument sets how many models this search will fit?
argmin = fmin(
fn=objective,
space=search_space,
algo=algo,
? =16)max_evals caps the number of points tested in the space. n_trials is Optuna's name for the same idea, max_queue_len controls how many settings are generated ahead of time, and parallelism belongs to SparkTrials.
Source: docs.databricks.comfmin() returns argmin, the best setting it found. There is a catch with categorical parameters: when you use hp.choice(), Hyperopt returns the index into the choice list, not the value, and that index is also what gets logged to MLflow. Call hyperopt.space_eval() to turn it back into the real value. Hyperopt also supports conditional hyperparameters, which apply only to some branches of a choice. For long-running models, Databricks suggests starting with small datasets and many hyperparameters, then using MLflow to see which hyperparameters can be fixed before tuning at scale.
Checkpoint 4 of 7· Exam question
A team is tuning a logistic regression model with only two hyperparameters: regularization strength `C` (4 candidate values) and penalty type (2 candidate values). For a compliance review, they need a documented guarantee that the best-performing combination among those specified values is found. Which search approach satisfies this requirement?
Correct answer: A — Use `GridSearchCV` over the 4 x 2 grid, since it exhaustively evaluates all eight combinations and is guaranteed to identify the best-performing one among the specified values.
- A. This is correct: `GridSearchCV` fits a model for every combination in the specified grid, so all eight combinations are evaluated and the best-performing one among them is guaranteed to be found.
- B. Random sampling with `n_iter=8` draws eight candidates independently and can sample the same combination more than once, so it does not guarantee that all eight distinct grid combinations are actually visited.
- C. The Tree-structured Parzen Estimator is a heuristic that prioritizes promising regions based on prior trials; it does not exhaustively evaluate every combination and offers no guarantee of finding the true best value among the specified options.
- D. Sampling `C` from a continuous distribution produces arbitrary real-valued draws and is very unlikely to land exactly on the four specified grid values, so it does not guarantee coverage of the discrete grid.
Sources3
4.Trials vs SparkTrials: where the trials run
The trials argument decides where each trial runs. SparkTrials distributes trials for single-machine models such as scikit-learn: the driver generates the trials and the Spark workers evaluate them. To switch it on, you add one more argument to fmin(), and nothing else in your Hyperopt code has to change:
spark_trials = SparkTrials()
with mlflow.start_run():
argmin = fmin(
fn=objective,
space=search_space,
algo=algo,
max_evals=16,
trials=spark_trials)For algorithms that already distribute their own training, such as MLlib or Horovod, use the default Trials class instead. Each trial then runs from the driver, so the algorithm can use the full cluster. SparkTrials also has limits. Do not use it on autoscaling clusters, because it fixes its parallelism when the run starts. For very short trials, Hyperopt and Spark overhead can cancel out most of the speedup.
Checkpoint 5 of 7· Check yourself
You are tuning a Spark MLlib model with Hyperopt. Which trials setting should you use?
MLlib already distributes training across the cluster, so trials run from the driver using the default Trials class. SparkTrials is for single-machine libraries such as scikit-learn.
“Use Trials when you call distributed training algorithms such as MLlib methods or Horovod in the objective function.”Source: docs.databricks.com
Checkpoint 6 of 7· Exam question
A data scientist runs `GridSearchCV` with `param_grid = {'max_depth': [3, 5, 7], 'min_samples_leaf': [1, 2, 5, 10]}` and `cv=5` on a random forest classifier. Assuming no combination is skipped, how many individual models does this fit in total?
Correct answer: C — 60 model fits
- A. This equals the number of hyperparameter combinations (3 x 4 = 12) but omits the 5-fold cross-validation refit for each combination, so it undercounts the total fits.
- B. This comes from adding the counts of the two parameter lists (3 + 4 = 7) and then multiplying by 5, which mixes up how the grid combines parameter values and does not reflect the actual combinatorics.
- C. This is correct: there are 3 x 4 = 12 hyperparameter combinations, and `cv=5` fits a separate model for each combination on every fold, giving 12 x 5 = 60 total model fits.
- D. This applies the 5-fold multiplier twice (12 x 5 x 5 = 300), which double-counts the cross-validation step and overstates the true number of fits.
Sources3
5.The same search in Optuna
Optuna, which Databricks recommends now, follows the same pattern with different names. The search space is defined inside the objective by calling suggest_* methods on a trial object. Because the space is written in code, it can be dynamic: in the example below, which hyperparameters exist depends on which model was chosen.
def objective(trial):
# Invoke suggest methods of a Trial object to generate hyperparameters.
regressor_name = trial.suggest_categorical('classifier', ['SVR', 'RandomForest'])
if regressor_name == 'SVR':
svr_c = trial.suggest_float('svr_c', 1e-10, 1e10, log=True)
regressor_obj = sklearn.svm.SVR(C=svr_c)
else:
rf_max_depth = trial.suggest_int('rf_max_depth', 2, 32)
regressor_obj = sklearn.ensemble.RandomForestRegressor(max_depth=rf_max_depth)You run the search with mlflow_study.optimize(objective, n_trials=8, n_jobs=4) on an MlflowSparkStudy, which runs trials in parallel on PySpark executors. Its default sampler is TPE, and its default pruner is MedianPruner, which stops unpromising trials early. Optuna minimizes by default too, which is why the Databricks getting-started example returns -roc_auc.
Checkpoint 7 of 7· Check yourself
In Optuna, where is the hyperparameter search space defined?
Optuna builds the space as the objective runs, using calls such as suggest_float and suggest_int. This is what allows conditional, dynamic spaces.
“Within the objective function, define the hyperparameter search space.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.SparkTrials is the right choice whenever you tune on a Spark cluster, including for MLlib models.Why is that wrong?
SparkTrials is for single-machine algorithms. Distributed algorithms such as MLlib already parallelize their training and should use the default Trials class.
Covered in Trials vs SparkTrials: where the trials run
2.The value fmin returns for an hp.choice parameter is the chosen option itself.Why is that wrong?
For hp.choice, Hyperopt returns and logs the index into the list. Use hyperopt.space_eval() to recover the actual value.
Covered in Search space, algorithm and the fmin() call
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-conceptsOfficial docs
“Hyperopt is not included in Databricks Runtime for Machine Learning after 16.4 LTS ML.”
↩︎ The four-step Hyperopt workflow“This function can return the loss as a scalar value or in a dictionary”
↩︎ The objective function returns a loss“Use SparkTrials when you call single-machine algorithms such as scikit-learn methods in the objective function.”
↩︎ Exam trap 1“Use Trials when you call distributed training algorithms such as MLlib methods or Horovod in the objective function.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-spark-mlflow-integrationOfficial docs
“Define a search space over hyperparameters.”
↩︎ The four-step Hyperopt workflow“Run the tuning algorithm with Hyperopt fmin().”
↩︎ Checkpoint“Hyperopt tries to minimize the objective function.”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-best-practicesOfficial docs
“Use hyperopt.space_eval() to retrieve the parameter values.”
↩︎ Search space, algorithm and the fmin() call“Take advantage of Hyperopt support for conditional dimensions and hyperparameters.”
↩︎ Search space, algorithm and the fmin() call“Do not use SparkTrials on autoscaling clusters.”
↩︎ Trials vs SparkTrials: where the trials run“Both Hyperopt and Spark incur overhead that can dominate the trial duration for short trial runs (low tens of seconds).”
↩︎ Trials vs SparkTrials: where the trials run“When you use hp.choice(), Hyperopt returns the index of the choice list.”
↩︎ Exam trap 2 - 4.
“MlflowSparkStudy class enables launching parallel Optuna studies using PySpark executors.”
↩︎ The same search in Optuna“Within the objective function, define the hyperparameter search space.”
↩︎ Checkpoint - 5.
“Negate the AUC because Optuna minimizes the objective by default”
↩︎ The same search in Optuna