CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 33/48

    Grid, random and Bayesian hyperparameter search compared

    Perform random or grid search or Bayesian search as a method for tuning hyperparameters.

    9 min read
    2.08% of exam
    7 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what one trial of a hyperparameter search is and how its cost is counted
    • Tell grid search, random search and Bayesian (TPE) search apart by how each picks the next setting
    • Choose a search strategy and a search-space shape for a tuning scenario
    • Explain the trade-off between parallelism and adaptivity in Bayesian search

    Key concept

    Adaptive vs non-adaptive search — Grid and random search decide which hyperparameter settings to try without looking at earlier scores. Bayesian search such as TPE uses the losses of earlier trials to decide which setting to try next.

    1.What a hyperparameter search actually does

    A hyperparameter search runs the same loop again and again. It picks a hyperparameter setting, trains a model with it, scores that model, and keeps the best one. Hyperopt calls each pass of the loop a trial, and a trial generally means fitting one model on one setting of hyperparameters. That makes the budget easy to count: if you ask for 16 evaluations, you pay for 16 model fits. In Hyperopt that budget is the max_evals argument, described as the number of hyperparameter settings to try, which is also the number of models to fit.

    The three methods differ in one place only: how the next setting is chosen. Grid search walks through a fixed list of combinations. Random search samples from the space. Bayesian search looks at the results so far before choosing. Training, scoring and the cost of each trial work the same way in all three. The rest of this page goes through the choice rule of each one.

    Checkpoint 1 of 5· Check yourself

    A Hyperopt run is configured with max_evals=16. Roughly how many models will it fit and evaluate?

    Sources1

    In grid search you list the candidate values for each hyperparameter, and the search evaluates the combinations of those lists. The scikit-learn cross-validation guide notes that the best parameters can be determined by grid search techniques. Its pipeline documentation shows GridSearchCV tuning a whole pipeline. In that example the grid covers which dimensionality-reduction step to use, which classifier to use, and the classifier's C value.

    A scikit-learn grid: three candidate values for reduce_dim, two classifiers and three values of C, passed to GridSearchCVpython
    >>> param_grid = dict(reduce_dim=['passthrough', PCA(5), PCA(10)],
    ...                   clf=[SVC(), LogisticRegression()],
    ...                   clf__C=[0.1, 10, 100])
    >>> grid_search = GridSearchCV(pipe, param_grid=param_grid)

    The Spark MLlib equivalents are CrossValidator and TrainValidationSplit. Older Databricks runtimes logged their hyperparameters and metrics to MLflow automatically, but that MLlib automated tracking is deprecated and is disabled by default from Databricks Runtime 10.4 LTS ML. A grid's main weakness is that it only tests the values you wrote down. It cannot try anything between C=10 and C=100.

    Checkpoint 2 of 5· Check yourself

    Using the grid above, a colleague adds a fourth candidate value to clf__C. How many combinations does the grid now hold?

    Sources23

    Random search does not use a list of values. You describe each hyperparameter as a distribution or a categorical choice, and the search draws settings from it. In Hyperopt, hyperopt.rand.suggest selects random search, which Databricks describes as a non-adaptive approach that samples over the search space. The space can mix categorical options, such as which algorithm to use, with numeric distributions such as uniform and log.

    A Ray Tune search space: continuous ranges (loguniform, uniform) sit next to categorical choicespython
        param_space={
            "lr": tune.loguniform(1e-5, 1e-3),
            "lora_r": tune.choice([8, 16, 32]),
            "lora_alpha_ratio": tune.choice([1, 2]),
            "lora_dropout": tune.uniform(0.0, 0.1),
            "weight_decay": tune.choice([0.0, 0.01]),
            "batch_size": tune.choice([4, 8]),
        },

    Because lr comes from a log-uniform range, two trials almost never share the same learning rate, so the search can land on values a grid would never list. Random sampling is also the default behaviour in Ray Tune. The Databricks example notes that if you want configurations to be *chosen* rather than sampled randomly, you pass TuneConfig a search_alg such as Optuna.

    Checkpoint 3 of 5· Check yourself

    What makes hyperopt.rand.suggest a non-adaptive search algorithm?

    Sources14

    Bayesian search keeps a record of which settings produced which losses and uses it to choose the next trial. In Hyperopt this is hyperopt.tpe.suggest, the Tree of Parzen Estimators. Databricks' best-practices page says Bayesian approaches can be much more efficient than grid search and random search, so with TPE you can explore more hyperparameters and wider ranges. The same page adds that using domain knowledge to narrow the search domain still improves the results. TPE is not limited to Hyperopt either: Optuna's MlflowSparkStudy uses optuna.samplers.TPESampler as its default sampler.

    This is the trade-off between parallelism and adaptivity. Running more trials at once finishes a fixed budget sooner. Running fewer at once gives each new proposal more evidence to work from. Random search has nothing to learn from earlier trials, so this trade-off does not apply to it. Also, do not expect the loss to fall with every trial. Hyperopt's search is stochastic, so the loss usually does not decrease monotonically, although these methods often find the best hyperparameters faster than other methods.

    How each search strategy chooses the next hyperparameter setting
    StrategyHow the next setting is chosenExample in the sources
    Grid searchNext combination from the fixed lists you suppliedGridSearchCV(pipe, param_grid=param_grid)
    Random searchSampled from the search space; non-adaptivehyperopt.rand.suggest
    Bayesian (TPE)Iteratively and adaptively, based on past resultshyperopt.tpe.suggest; optuna.samplers.TPESampler

    Checkpoint 4 of 5· Match them up

    Match each search algorithm to its behaviour

    Tap a term, then the definition that fits it.

    Checkpoint 5 of 5· Exam question

    A data scientist is tuning a gradient boosting model with six hyperparameters, several of which are continuous (for example, learning rate and subsample ratio). Compute time is capped at a fixed budget of 40 model fits, and preliminary analysis suggests only two or three of the six hyperparameters meaningfully affect performance. Which search strategy should the data scientist use to stay within budget while tuning effectively?

    Sources56

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.With TPE, running as many trials in parallel as possible always gives the best result.Why is that wrong?

      TPE adapts using past results. For a fixed max_evals, more parallelism is faster, but lower parallelism may give better results because each proposal has more completed trials to learn from.

      Covered in Bayesian search: letting past trials pick the next one

    2. 2.If the loss goes up between consecutive Hyperopt trials, the search is broken.Why is that wrong?

      Hyperopt's algorithms are stochastic, so the loss is not expected to fall with every run.

      Covered in Bayesian search: letting past trials pick the next one

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “In Hyperopt, a trial generally corresponds to fitting one model on one setting of hyperparameters.”
      ↩︎ What a hyperparameter search actually does
      “Number of hyperparameter settings to try (the number of models to fit).”
      ↩︎ What a hyperparameter search actually does
      “You can choose a categorical option such as algorithm, or probabilistic distribution for numeric values such as uniform and log.”
      ↩︎ Random search: sampling from distributions
      “Because Hyperopt proposes new trials based on past results, there is a trade-off between parallelism and adaptivity.”
      ↩︎ Exam trap 1
      “lower parallelism may lead to better results since each iteration has access to more past results”
      ↩︎ Prediction
      “Most commonly used are hyperopt.rand.suggest for Random Search and hyperopt.tpe.suggest for TPE.”
      ↩︎ Checkpoint
    2. 2.
      “MLlib automated MLflow tracking is deprecated and disabled by default on clusters that run Databricks Runtime 10.4 LTS ML and above.”
      ↩︎ Grid search: every combination of a fixed list
      “when you run tuning code that uses CrossValidator or TrainValidationSplit, hyperparameters and evaluation metrics are automatically logged in MLflow.”
      ↩︎ Grid search: every combination of a fixed list
    3. 3.
      “The best parameters can be determined by grid search techniques.”
      ↩︎ Grid search: every combination of a fixed list
    4. 4.
      “To choose configurations instead of sampling them randomly, pass TuneConfig a search_alg such as Optuna.”
      ↩︎ Random search: sampling from distributions
    5. 5.
      “Bayesian approaches can be much more efficient than grid search and random search.”
      ↩︎ Bayesian search: letting past trials pick the next one
      “Using domain knowledge to restrict the search domain can optimize tuning and produce better results.”
      ↩︎ Bayesian search: letting past trials pick the next one
      “However, these methods often find the best hyperparameters more quickly than other methods.”
      ↩︎ Bayesian search: letting past trials pick the next one
      “Because Hyperopt uses stochastic search algorithms, the loss usually does not decrease monotonically with each run.”
      ↩︎ Exam trap 2

    Also cited

    Continue to page 2 of 2

    Running a hyperparameter search with Hyperopt fmin and Optuna

    Spotted a mistake, or was something unclear? Tell us.