What you will be able to do
- Explain what one trial of a hyperparameter search is and how its cost is counted
- Tell grid search, random search and Bayesian (TPE) search apart by how each picks the next setting
- Choose a search strategy and a search-space shape for a tuning scenario
- Explain the trade-off between parallelism and adaptivity in Bayesian search
Key concept
Adaptive vs non-adaptive search — Grid and random search decide which hyperparameter settings to try without looking at earlier scores. Bayesian search such as TPE uses the losses of earlier trials to decide which setting to try next.
1.What a hyperparameter search actually does
A hyperparameter search runs the same loop again and again. It picks a hyperparameter setting, trains a model with it, scores that model, and keeps the best one. Hyperopt calls each pass of the loop a trial, and a trial generally means fitting one model on one setting of hyperparameters. That makes the budget easy to count: if you ask for 16 evaluations, you pay for 16 model fits. In Hyperopt that budget is the max_evals argument, described as the number of hyperparameter settings to try, which is also the number of models to fit.
The three methods differ in one place only: how the next setting is chosen. Grid search walks through a fixed list of combinations. Random search samples from the space. Bayesian search looks at the results so far before choosing. Training, scoring and the cost of each trial work the same way in all three. The rest of this page goes through the choice rule of each one.
Checkpoint 1 of 5· Check yourself
A Hyperopt run is configured with max_evals=16. Roughly how many models will it fit and evaluate?
max_evals is the maximum number of points in the space to test, and each point is one fitted model, so 16 evaluations means about 16 model fits.
“Set max_evals to the maximum number of points in hyperparameter space to test, that is, the maximum number of models to fit and evaluate.”Source: docs.databricks.com
Sources1
2.Grid search: every combination of a fixed list
In grid search you list the candidate values for each hyperparameter, and the search evaluates the combinations of those lists. The scikit-learn cross-validation guide notes that the best parameters can be determined by grid search techniques. Its pipeline documentation shows GridSearchCV tuning a whole pipeline. In that example the grid covers which dimensionality-reduction step to use, which classifier to use, and the classifier's C value.
>>> param_grid = dict(reduce_dim=['passthrough', PCA(5), PCA(10)],
... clf=[SVC(), LogisticRegression()],
... clf__C=[0.1, 10, 100])
>>> grid_search = GridSearchCV(pipe, param_grid=param_grid)3 × 2 × 3 = 18. Each new list multiplies the total, so a grid grows quickly as you add hyperparameters or values. Every combination also costs a full model fit.
The Spark MLlib equivalents are CrossValidator and TrainValidationSplit. Older Databricks runtimes logged their hyperparameters and metrics to MLflow automatically, but that MLlib automated tracking is deprecated and is disabled by default from Databricks Runtime 10.4 LTS ML. A grid's main weakness is that it only tests the values you wrote down. It cannot try anything between C=10 and C=100.
Checkpoint 2 of 5· Check yourself
Using the grid above, a colleague adds a fourth candidate value to clf__C. How many combinations does the grid now hold?
The grid is the product of its lists: 3 reduce_dim options × 2 classifiers × 4 C values = 24. Adding a value multiplies the total; it does not add one.
“The best parameters can be determined by grid search techniques.”Source: scikit-learn.org
3.Random search: sampling from distributions
Random search does not use a list of values. You describe each hyperparameter as a distribution or a categorical choice, and the search draws settings from it. In Hyperopt, hyperopt.rand.suggest selects random search, which Databricks describes as a non-adaptive approach that samples over the search space. The space can mix categorical options, such as which algorithm to use, with numeric distributions such as uniform and log.
param_space={
"lr": tune.loguniform(1e-5, 1e-3),
"lora_r": tune.choice([8, 16, 32]),
"lora_alpha_ratio": tune.choice([1, 2]),
"lora_dropout": tune.uniform(0.0, 0.1),
"weight_decay": tune.choice([0.0, 0.01]),
"batch_size": tune.choice([4, 8]),
},Because lr comes from a log-uniform range, two trials almost never share the same learning rate, so the search can land on values a grid would never list. Random sampling is also the default behaviour in Ray Tune. The Databricks example notes that if you want configurations to be *chosen* rather than sampled randomly, you pass TuneConfig a search_alg such as Optuna.
Checkpoint 3 of 5· Check yourself
What makes hyperopt.rand.suggest a non-adaptive search algorithm?
Random search draws each setting from the defined space independently. Earlier losses do not influence where it samples next.
“hyperopt.rand.suggest: Random search, a non-adaptive approach that samples over the search space”Source: docs.databricks.com
4.Bayesian search: letting past trials pick the next one
Bayesian search keeps a record of which settings produced which losses and uses it to choose the next trial. In Hyperopt this is hyperopt.tpe.suggest, the Tree of Parzen Estimators. Databricks' best-practices page says Bayesian approaches can be much more efficient than grid search and random search, so with TPE you can explore more hyperparameters and wider ranges. The same page adds that using domain knowledge to narrow the search domain still improves the results. TPE is not limited to Hyperopt either: Optuna's MlflowSparkStudy uses optuna.samplers.TPESampler as its default sampler.
This is the trade-off between parallelism and adaptivity. Running more trials at once finishes a fixed budget sooner. Running fewer at once gives each new proposal more evidence to work from. Random search has nothing to learn from earlier trials, so this trade-off does not apply to it. Also, do not expect the loss to fall with every trial. Hyperopt's search is stochastic, so the loss usually does not decrease monotonically, although these methods often find the best hyperparameters faster than other methods.
| Strategy | How the next setting is chosen | Example in the sources |
|---|---|---|
| Grid search | Next combination from the fixed lists you supplied | GridSearchCV(pipe, param_grid=param_grid) |
| Random search | Sampled from the search space; non-adaptive | hyperopt.rand.suggest |
| Bayesian (TPE) | Iteratively and adaptively, based on past results | hyperopt.tpe.suggest; optuna.samplers.TPESampler |
Checkpoint 4 of 5· Match them up
Match each search algorithm to its behaviour
Tap a term, then the definition that fits it.
TPE is the Bayesian, adaptive option, rand.suggest is non-adaptive random sampling, and GridSearchCV works through an explicit grid of values.
“Most commonly used are hyperopt.rand.suggest for Random Search and hyperopt.tpe.suggest for TPE.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A data scientist is tuning a gradient boosting model with six hyperparameters, several of which are continuous (for example, learning rate and subsample ratio). Compute time is capped at a fixed budget of 40 model fits, and preliminary analysis suggests only two or three of the six hyperparameters meaningfully affect performance. Which search strategy should the data scientist use to stay within budget while tuning effectively?
Correct answer: A — Use `RandomizedSearchCV` with `n_iter=40`, since the evaluation budget stays fixed regardless of dimensionality and hyperparameters that do not influence performance do not reduce sampling efficiency.
- A. This is correct: with a fixed `n_iter`, the total number of fits never grows with the number of hyperparameters, and because random sampling does not need to cover every combination, hyperparameters with little real effect on the score do not cost extra evaluations.
- B. A coarse grid of two values across six hyperparameters produces 2^6 = 64 combinations, which already exceeds the 40-fit budget, so grid search cannot guarantee staying within the limit as dimensionality grows.
- C. `GridSearchCV` requires an explicit list of discrete values for each parameter in `param_grid`; it has no mechanism to accept or discretize a continuous distribution object, so this configuration is not valid.
- D. The `n_iter` parameter of `RandomizedSearchCV` defaults to a fixed value and is set explicitly by the user; it does not automatically scale with the number of hyperparameters being tuned.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.With TPE, running as many trials in parallel as possible always gives the best result.Why is that wrong?
TPE adapts using past results. For a fixed max_evals, more parallelism is faster, but lower parallelism may give better results because each proposal has more completed trials to learn from.
Covered in Bayesian search: letting past trials pick the next one
2.If the loss goes up between consecutive Hyperopt trials, the search is broken.Why is that wrong?
Hyperopt's algorithms are stochastic, so the loss is not expected to fall with every run.
Covered in Bayesian search: letting past trials pick the next one
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-conceptsOfficial docs
“In Hyperopt, a trial generally corresponds to fitting one model on one setting of hyperparameters.”
↩︎ What a hyperparameter search actually does“Number of hyperparameter settings to try (the number of models to fit).”
↩︎ What a hyperparameter search actually does“You can choose a categorical option such as algorithm, or probabilistic distribution for numeric values such as uniform and log.”
↩︎ Random search: sampling from distributions“Because Hyperopt proposes new trials based on past results, there is a trade-off between parallelism and adaptivity.”
↩︎ Exam trap 1“lower parallelism may lead to better results since each iteration has access to more past results”
↩︎ Prediction“Most commonly used are hyperopt.rand.suggest for Random Search and hyperopt.tpe.suggest for TPE.”
↩︎ Checkpoint - 2.
“MLlib automated MLflow tracking is deprecated and disabled by default on clusters that run Databricks Runtime 10.4 LTS ML and above.”
↩︎ Grid search: every combination of a fixed list“when you run tuning code that uses CrossValidator or TrainValidationSplit, hyperparameters and evaluation metrics are automatically logged in MLflow.”
↩︎ Grid search: every combination of a fixed list - 3.https://scikit-learn.org/stable/modules/cross_validation.htmlSecondary source
“The best parameters can be determined by grid search techniques.”
↩︎ Grid search: every combination of a fixed list - 4.https://docs.databricks.com/aws/en/machine-learning/ai-runtime/cli/examples/ray-tune-loraOfficial docs
“To choose configurations instead of sampling them randomly, pass TuneConfig a search_alg such as Optuna.”
↩︎ Random search: sampling from distributions - 5.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-best-practicesOfficial docs
“Bayesian approaches can be much more efficient than grid search and random search.”
↩︎ Bayesian search: letting past trials pick the next one“Using domain knowledge to restrict the search domain can optimize tuning and produce better results.”
↩︎ Bayesian search: letting past trials pick the next one“However, these methods often find the best hyperparameters more quickly than other methods.”
↩︎ Bayesian search: letting past trials pick the next one“Because Hyperopt uses stochastic search algorithms, the loss usually does not decrease monotonically with each run.”
↩︎ Exam trap 2 - 6.
“optuna.samplers.TPESampler is used as the default.”
↩︎ Bayesian search: letting past trials pick the next one
Also cited
- https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-spark-mlflow-integrationOfficial docs
“a Bayesian approach which iteratively and adaptively selects new hyperparameter settings to explore based on past results”
↩︎ Key concept“Set max_evals to the maximum number of points in hyperparameter space to test, that is, the maximum number of models to fit and evaluate.”
↩︎ Checkpoint“hyperopt.rand.suggest: Random search, a non-adaptive approach that samples over the search space”
↩︎ Checkpoint