What you will be able to do
- Choose the right trials argument for fmin(): Trials for distributed training algorithms, SparkTrials for single-machine ones
- Use max_queue_len and early_stop_fn to control how fmin() generates and stops trials
- Wrap an fmin() call in an MLflow run so each tuning run is logged separately
- Read a tuning run correctly: NaN losses, loss curves that go up and down, and when to narrow the space
1.fmin() arguments beyond fn, space and algo
A minimal fmin() call takes an objective function (fn), a search space (space), a search algorithm (algo) and a budget (max_evals, the number of hyperparameter settings to try). The function also takes optional arguments that control how trials are generated, where they run and when the search stops. On the exam they are usually tested by asking which one solves a given problem.
| Argument | What it controls | Notes from the docs |
|---|---|---|
| max_queue_len | How many hyperparameter settings Hyperopt generates ahead of time | Default 1. Raising it can help because TPE generation can take time, but generally keep it no larger than the SparkTrials parallelism setting |
| early_stop_fn | Whether fmin() stops before max_evals is reached | Default None. Takes Trials and *args, and returns a bool plus *args. The returned *args are passed into the next call |
| trials | Which object records and runs the trials | A Trials or SparkTrials object |
early_stop_fn answers the prediction above. It lets the run end early once it decides further trials are not worth running, so max_evals acts as an upper limit. The *args it returns are state that carries over from one call to the next, which lets it track something like how many trials have passed without improvement. There is one caveat when trials are distributed with SparkTrials: the early stopping function is polled, so it is not guaranteed to run after every trial.
Checkpoint 1 of 5· Match them up
Match each fmin() argument to what it controls.
Tap a term, then the definition that fits it.
max_evals sets the budget, max_queue_len sets how far ahead settings are generated, early_stop_fn can end the run early, and trials picks how and where trials run.
“An optional early stopping function to determine if fmin should stop before max_evals is reached.”Source: docs.databricks.com
Sources1
2.Choosing the trials argument: Trials or SparkTrials
The trials choice depends on what the objective function trains. If it calls a single-machine library such as scikit-learn, pass a SparkTrials object. Databricks built it so that separate trials can run at the same time on Spark workers. If it calls a distributed training algorithm, such as Spark MLlib or Horovod, use the default Trials class. Each trial then runs from the driver node and has the full cluster available, so the algorithm can start its own distributed training.
The choice also affects logging. Databricks does not support automatic MLflow logging with the Trials class, so with distributed algorithms you have to log trials to MLflow yourself. With SparkTrials, Databricks recommends wrapping the fmin() call in with mlflow.start_run():, which gives each call its own main MLflow run. Moving the scikit-learn SVC example to SparkTrials adds just one argument:
Checkpoint 2 of 5· Fill the gap
Complete the call so the scikit-learn trials are distributed to Spark workers.
spark_trials = SparkTrials()
with mlflow.start_run():
argmin = fmin(
fn=objective,
space=search_space,
algo=algo,
max_evals=16,
trials= ? )The SparkTrials instance goes in the trials argument. The rest of the fmin() call is the same as in a single-machine run.
Source: docs.databricks.comCheckpoint 3 of 5· Exam question
A team is defining the search space for `fmin` to tune a gradient-boosted tree model. `n_estimators` must always be a whole number of trees, while `learning_rate` should be sampled on a scale where small values are explored proportionally more finely than large ones. Which pairing of Hyperopt space functions matches these two requirements?
Correct answer: A — Use `hp.quniform` for `n_estimators` so the sampled value rounds to an integer step size, and use `hp.loguniform` for `learning_rate` so values are drawn evenly across orders of magnitude.
- A. This is correct: `hp.quniform` rounds continuous samples to a chosen step so `n_estimators` stays an integer count of trees, and `hp.loguniform` samples in log space so small learning rates get proportionally as much exploration as large ones.
- B. This is incorrect: plain `hp.uniform` returns a float, so `n_estimators` would not reliably land on whole tree counts, and applying the same linear scheme to `learning_rate` ignores that small rates need finer relative exploration.
- C. This is incorrect: `hp.choice` restricts `n_estimators` to a hand-picked list rather than letting the search explore a range, and `hp.normal` concentrates `learning_rate` samples near zero instead of spanning the useful range evenly.
- D. This is incorrect: `hp.randint` forces the range to start at zero and offers no step control for `n_estimators`, and a linear `hp.uniform` for `learning_rate` under-samples the small-value region that usually matters most.
3.Reading and refining a tuning run
When the run finishes, you usually look at the loss of each trial, for example by comparing runs in MLflow. Two things in that view can look like bugs but usually aren't.
The first is a reported loss of NaN. This usually means the objective function returned NaN for that hyperparameter setting. It does not affect the other trials and can safely be ignored. If you want to avoid it, adjust the search space or change the objective function.
The second is a loss that goes up and down from trial to trial. Hyperopt's search algorithms are stochastic, so the loss usually does not fall steadily with each run. Databricks notes that these methods often still find the best hyperparameters faster than other methods.
Checkpoint 4 of 5· Check yourself
After 40 trials with tpe.suggest, the loss rises and falls from trial to trial instead of dropping steadily, and one trial reports NaN. What is the best reading?
A loss that is not monotonic is expected from stochastic search. A NaN loss comes from the objective's return value for one setting and does not affect other runs.
“Because Hyperopt uses stochastic search algorithms, the loss usually does not decrease monotonically with each run.”Source: docs.databricks.com
Results also help you plan the next run. For models that take a long time to train, Databricks suggests starting with small datasets and many hyperparameters. Use MLflow to find the best-performing models and work out which hyperparameters can be fixed, then tune a smaller space at full scale. Keep the trial length in mind too. Hyperopt and Spark both add overhead, and for trials of only a few tens of seconds that overhead can wipe out most or all of the speedup from distributing them.
Checkpoint 5 of 5· Exam question
Within Hyperopt, what is the key behavioral difference between passing `algo=hyperopt.tpe.suggest` versus `algo=hyperopt.rand.suggest` to `fmin`?
Correct answer: A — `tpe.suggest` uses results from earlier trials to model which regions of the search space look promising, while `rand.suggest` draws each configuration independently of prior outcomes.
- A. This is correct: the Tree-structured Parzen Estimator behind `tpe.suggest` builds a probabilistic model from previously evaluated trials to propose promising next points, whereas random search samples each new point independently of history.
- B. This is incorrect: neither search algorithm requires a specific `Trials` implementation; both `tpe.suggest` and `rand.suggest` can be paired with either plain `Trials` or `SparkTrials`.
- C. This is incorrect: both algorithms can sample from any combination of `hp.choice`, `hp.uniform`, and other space definitions; neither is restricted to discrete-only or continuous-only parameters.
- D. This is incorrect: no Hyperopt search algorithm guarantees finding the global optimum, and random search does not have a special limitation that caps it at a local optimum; both are heuristic searches over a finite evaluation budget.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Always pass SparkTrials to fmin() on Databricks, even when the objective trains a Spark MLlib or Horovod model.Why is that wrong?
SparkTrials is for single-machine algorithms such as scikit-learn. Distributed algorithms should use the default Trials class, so that each trial runs from the driver and the algorithm can distribute its own training.
Covered in Choosing the trials argument: Trials or SparkTrials
2.If the loss does not fall with each Hyperopt trial, the tuning is misconfigured.Why is that wrong?
Hyperopt's search algorithms are stochastic, so a loss that rises and falls between trials is expected. A NaN loss means the objective returned NaN for that one setting and does not affect other runs.
Covered in Reading and refining a tuning run
Practise it for real
Tune the regularization parameter C of a scikit-learn SVC on the Iris dataset with fmin(), first on a single machine and then with SparkTrials
1.Load Iris with load_iris(), and define objective(C) so it builds SVC(C=C), computes cross_val_score(clf, X, y).mean(), and returns {'loss': -accuracy, 'status': STATUS_OK}
Why: fmin() minimizes the loss, so the accuracy has to be negated
You should see: Calling objective(1.0) returns a dictionary with a negative 'loss' value
2.Define search_space = hp.lognormal('C', 0, 1.0) and set algo=tpe.suggest
Why: These become the space and algo arguments. TPE uses earlier results to pick the next settings
You should see: Both objects are defined and ready to pass to fmin()
3.Run argmin = fmin(fn=objective, space=search_space, algo=algo, max_evals=16) and print argmin
Why: This runs 16 trials on a single machine, fitting 16 models
You should see: A dictionary with the best value found for C
4.Create spark_trials = SparkTrials(), then rerun the same fmin() call inside with mlflow.start_run(): with trials=spark_trials added
Why: SparkTrials distributes the single-machine scikit-learn trials, and the MLflow run groups all the trials of this call
You should see: A new MLflow run with one child run per trial, plus a best C value
Stuck? Get a nudge
If the distributed run takes longer than the single-machine one, check how long each trial is. The objective here is fast, so Spark job startup costs more than the trials themselves.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-conceptsOfficial docs
“it can be helpful to increase this beyond the default value of 1, but generally no larger than the SparkTrials setting parallelism.”
↩︎ fmin() arguments beyond fn, space and algo“When using SparkTrials, the early stopping function is not guaranteed to run after every trial, and is instead polled.”
↩︎ fmin() arguments beyond fn, space and algo“Use SparkTrials when you call single-machine algorithms such as scikit-learn methods in the objective function.”
↩︎ Choosing the trials argument: Trials or SparkTrials“When calling fmin(), Databricks recommends active MLflow run management; that is, wrap the call to fmin() inside a with mlflow.start_run(): statement.”
↩︎ Choosing the trials argument: Trials or SparkTrials“Use Trials when you call distributed training algorithms such as MLlib methods or Horovod in the objective function.”
↩︎ Exam trap 1“An optional early stopping function to determine if fmin should stop before max_evals is reached.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-distributed-mlOfficial docs
“Databricks does not support automatic logging to MLflow with the Trials class.”
↩︎ Choosing the trials argument: Trials or SparkTrials“Hyperopt evaluates each trial on the driver node so that the ML algorithm itself can initiate distributed training.”
↩︎ Choosing the trials argument: Trials or SparkTrials“do not pass a trials argument to fmin(), and specifically, do not use the SparkTrials class.”
↩︎ Prediction - 3.https://docs.databricks.com/aws/en/machine-learning/automl-hyperparam-tuning/hyperopt-best-practicesOfficial docs
“This does not affect other runs and you can safely ignore it.”
↩︎ Reading and refining a tuning run“start experimenting with small datasets and many hyperparameters”
↩︎ Reading and refining a tuning run“Both Hyperopt and Spark incur overhead that can dominate the trial duration for short trial runs (low tens of seconds).”
↩︎ Reading and refining a tuning run“A reported loss of NaN (not a number) usually means the objective function passed to fmin() returned NaN.”
↩︎ Exam trap 2“Because Hyperopt uses stochastic search algorithms, the loss usually does not decrease monotonically with each run.”
↩︎ Checkpoint