CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 32/48

    Hyperopt fmin: Objective Function, Search Space and Search Algorithm

    Use Hyperopt's fmin operation to tune a model's hyperparameters

    9 min read
    2.08% of exam
    3 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • List the four steps of a Hyperopt workflow in the order they happen
    • Write an objective function that returns a loss fmin() can minimize, including negating a metric where higher is better
    • Define a search space with hp expressions and recover real values from hp.choice indexes with space_eval()
    • Choose between tpe.suggest and rand.suggest, set max_evals, and read the result fmin() returns

    Key concept

    fmin() minimizes a loss — fmin() calls your objective function again and again, each time with a hyperparameter setting drawn from the search space, and keeps the setting that gave the lowest loss. Any metric where higher is better, such as accuracy, has to be turned into a loss before fmin() can work with it.

    1.The four steps of a Hyperopt workflow

    Hyperopt tunes hyperparameters with one function, fmin(). You give it a function to minimize and a space of hyperparameter values to search. It tries settings from that space and returns the best one it found. The Databricks example tunes the regularization parameter C of a scikit-learn support vector classifier on the Iris dataset, and this lesson uses that example throughout.

    The documentation breaks the work into four steps. First, define a function to minimize. Second, define a search space over the hyperparameters. Third, select a search algorithm. Fourth, run the tuning with fmin(). Each of the first three steps produces one argument for the fourth: the function becomes fn, the space becomes space and the algorithm becomes algo. That is why fmin() comes last.

    Checkpoint 1 of 5· Put it in order

    Put the steps of a Hyperopt workflow in the order the Databricks documentation gives them.

    1. 1.Run the tuning algorithm with Hyperopt fmin()
    2. 2.Define a function to minimize
    3. 3.Select a search algorithm
    4. 4.Define a search space over hyperparameters

    Sources1

    2.Writing the objective function (fn)

    The documentation says most of the code in a Hyperopt workflow lives in the objective function. Hyperopt calls it once per trial, each time with values drawn from your search space. Inside, you typically train a model with those values, score it, and return a loss. The loss can be a plain scalar or a dictionary. In the Databricks example, the function builds an SVC with the proposed C and scores it with the mean cross-validated accuracy from cross_val_score.

    The objective function from the Databricks example: train, score with cross-validation, and return the negated accuracy as the losspython
    def objective(C):
        # Create a support vector classifier model
        clf = SVC(C=C)
    
        # Use the cross-validation accuracy to compare the models' performance
        accuracy = cross_val_score(clf, X, y).mean()
    
        # Hyperopt tries to minimize the objective function. A higher accuracy value means a better model, so you must return the negative accuracy.
        return {'loss': -accuracy, 'status': STATUS_OK}

    This example uses the dictionary form. The 'loss' key holds the number Hyperopt minimizes, and 'status': STATUS_OK marks the trial as successful. STATUS_OK is imported from hyperopt along with fmin, tpe and hp. The negation is the detail exams test. If you return accuracy unchanged, fmin() will look for the least accurate model.

    Checkpoint 2 of 5· Fill the gap

    Complete the return statement so that fmin() finds the most accurate model.

    return {'loss':  ? , 'status': STATUS_OK}

    Checkpoint 3 of 5· Exam question

    A data scientist is tuning a scikit-learn `RandomForestClassifier` with Hyperopt's `fmin` and wants the search to converge on the configuration with the highest cross-validated accuracy. Inside the objective function passed as `fn`, which return value correctly points `fmin` toward that configuration?

    Sources12

    3.Defining the search space (space)

    The space argument defines where Hyperopt looks. A dimension can be categorical, such as a choice between algorithms, or a probability distribution over numeric values. The documentation names uniform and log distributions as examples. The SVC example has a single dimension: C, drawn from a lognormal distribution and labelled 'C'.

    A one-dimensional search space for the SVC regularization parameter Cpython
    search_space = hp.lognormal('C', 0, 1.0)

    Categorical dimensions defined with hp.choice() have a catch. Hyperopt returns the position of the chosen item in the list, not the item itself, and that position is also what gets logged to MLflow. To turn the result back into real parameter values, pass it through hyperopt.space_eval().

    Search spaces can also be conditional. Say you are comparing several flavors of gradient descent. You don't have to limit the space to the hyperparameters they all share: Hyperopt can include hyperparameters that apply only to some of the flavors. Size the space with care. Domain knowledge that narrows the ranges usually makes tuning faster and gives better results.

    Checkpoint 4 of 5· Check yourself

    A search space contains hp.choice('model', ['svm', 'rf', 'lr']). The tuned result and the MLflow log show model = 1. What does that mean, and how do you get the actual value?

    Sources13

    4.Choosing the algorithm and running fmin()

    The two search algorithms you usually pass to fmin()'s algo argument
    algo valueApproachHow it picks the next setting
    hyperopt.tpe.suggestTree of Parzen Estimators (Bayesian)Iteratively and adaptively, based on the results of earlier trials
    hyperopt.rand.suggestRandom search (non-adaptive)Samples over the search space without using earlier results

    TPE uses what earlier trials found, and Databricks notes that Bayesian approaches can be much more efficient than grid search and random search. With TPE you can therefore afford more hyperparameters and wider ranges. The example sets algo=tpe.suggest.

    The last required decision is max_evals. It is the number of hyperparameter settings to try, which is also the number of models to fit and evaluate. It does not count epochs or iterations inside one model. With all four pieces ready, the run is a single call:

    Running the tuning: 16 settings of C, each a fitted and cross-validated SVCpython
    argmin = fmin(
      fn=objective,
      space=search_space,
      algo=algo,
      max_evals=16)
    
    # Print the best value found for C
    print("Best value found: ", argmin)

    fmin() returns the best hyperparameter values it found, here the best C. If the space used hp.choice(), those entries are indexes, so apply space_eval() before using the values to train a final model.

    Checkpoint 5 of 5· Check yourself

    A data scientist calls fmin() with max_evals=50 and algo=tpe.suggest. What does max_evals=50 control?

    Sources23

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Return the accuracy (or AUC, or F1) directly as the loss, and fmin() will find the best model.Why is that wrong?

      fmin() always minimizes. A metric where higher is better has to be negated, as in {'loss': -accuracy, ...}, or the search will go after the worst model.

      Covered in Writing the objective function (fn)

    2. 2.For an hp.choice() dimension, the value that fmin() returns and MLflow logs is the selected option itself.Why is that wrong?

      It is the index of that option in the choice list. Use hyperopt.space_eval() to get the actual parameter values back.

      Covered in Defining the search space (space)

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “You use fmin() to execute a Hyperopt run.”
      ↩︎ The four steps of a Hyperopt workflow
      “Hyperopt is not included in Databricks Runtime for Machine Learning after 16.4 LTS ML.”
      ↩︎ The four steps of a Hyperopt workflow
      “This function can return the loss as a scalar value or in a dictionary (see Hyperopt docs for details).”
      ↩︎ Writing the objective function (fn)
      “You can choose a categorical option such as algorithm, or probabilistic distribution for numeric values such as uniform and log.”
      ↩︎ Defining the search space (space)
      “Number of hyperparameter settings to try (the number of models to fit).”
      ↩︎ Checkpoint
    2. 2.
      “Most of the code for a Hyperopt workflow is in the objective function.”
      ↩︎ Writing the objective function (fn)
      “hyperopt.rand.suggest: Random search, a non-adaptive approach that samples over the search space”
      ↩︎ Choosing the algorithm and running fmin()
      “the maximum number of models to fit and evaluate”
      ↩︎ Choosing the algorithm and running fmin()
      “Hyperopt tries to minimize the objective function.”
      ↩︎ Key concept
      “A higher accuracy value means a better model, so you must return the negative accuracy.”
      ↩︎ Exam trap 1
      “Define a function to minimize.”
      ↩︎ Checkpoint
      “A higher accuracy value means a better model, so you must return the negative accuracy.”
      ↩︎ Prediction
    3. 3.
      “Using domain knowledge to restrict the search domain can optimize tuning and produce better results.”
      ↩︎ Defining the search space (space)
      “Take advantage of Hyperopt support for conditional dimensions and hyperparameters.”
      ↩︎ Defining the search space (space)
      “Bayesian approaches can be much more efficient than grid search and random search.”
      ↩︎ Choosing the algorithm and running fmin()
      “Use hyperopt.space_eval() to retrieve the parameter values.”
      ↩︎ Exam trap 2
      “When you use hp.choice(), Hyperopt returns the index of the choice list.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Hyperopt fmin Control Arguments: trials, early_stop_fn and Reading Results

    Spotted a mistake, or was something unclear? Tell us.