CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 31/48

    Hyperparameter Tuning on Databricks: Cross-Validation, Hyperopt fmin, and SparkTrials

    Develop a training pipeline

    9 min read
    2.08% of exam
    6 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Compare k-fold cross-validation with a single validation split and count the models a tuning run fits
    • Write a Hyperopt objective, search space and fmin call, and choose between tpe.suggest and rand.suggest
    • Decide when to use SparkTrials versus Trials and how parallelism affects results

    1.Cross-validation versus a single validation split

    Tuning means comparing candidate settings, and each comparison needs a reliable score. k-fold cross-validation splits the data into k folds. Each fold takes one turn as the validation set while the model trains on the others, and the reported score is the average across those k rounds. A fixed validation split fits only one model per setting but holds back data. Cross-validation costs k fits per setting but uses all the data for validation. That gives the counting rule: models fitted = hyperparameter combinations × folds. On Spark ML, the equivalent tuning classes are CrossValidator and TrainValidationSplit. Their automatic MLflow tracking is deprecated in favour of mlflow.pyspark.ml.autolog().

    The simplest scikit-learn entry point is cross_val_score:

    Five-fold cross-validation of one estimator with cross_val_scorepython
    >>> from sklearn.model_selection import cross_val_score
    >>> clf = svm.SVC(kernel='linear', C=1, random_state=42)
    >>> scores = cross_val_score(clf, X, y, cv=5)

    Databricks' own Hyperopt example scores each candidate with cross_val_score inside the objective function, so every trial already averages over folds.

    Checkpoint 1 of 6· Exam question

    A training pipeline needs to preprocess a DataFrame containing both numeric columns (impute and scale) and categorical columns (one-hot encode) before fitting a `RandomForestClassifier`. Which approach correctly composes this into a single scikit-learn object that can be passed to `GridSearchCV`?

    Checkpoint 2 of 6· Check yourself

    Why might you accept the extra compute cost of k-fold cross-validation over a single fixed validation split?

    Sources123

    2.Tuning with Hyperopt fmin

    A Hyperopt workflow has four steps. You define a function to minimize, define a search space over hyperparameters, select a search algorithm, and run fmin(). Most of the code lives in the objective function, which usually trains a model and calculates its loss. Hyperopt always minimizes. If your score is "higher is better", such as accuracy, return its negative as the loss.

    The search space can mix categorical options with probability distributions such as uniform and log-based ones. The Databricks example tunes an SVM's regularization parameter C:

    A log-normal search space for the SVM parameter Cpython
    search_space = hp.lognormal('C', 0, 1.0)

    You have two main algorithms to choose from. hyperopt.rand.suggest is random search, which samples the space without adapting. hyperopt.tpe.suggest is the Tree of Parzen Estimators, a Bayesian method that picks new settings based on past results. Databricks notes that Bayesian approaches can be much more efficient than grid search and random search. That efficiency lets you explore more hyperparameters and wider ranges. max_evals caps how many settings are tried, which is also the number of models fitted. One quirk: with hp.choice(), Hyperopt returns and logs the *index* of the choice, so use hyperopt.space_eval() to recover the actual value.

    Checkpoint 3 of 6· Put it in order

    Put the steps of a Hyperopt workflow in order

    1. 1.Select a search algorithm
    2. 2.Define a function to minimize
    3. 3.Define a search space over hyperparameters
    4. 4.Run the tuning algorithm with Hyperopt fmin()

    Checkpoint 4 of 6· Fill the gap

    Which function scores each candidate C in the Databricks objective function?

    def objective(C):
        # Create a support vector classifier model
        clf = SVC(C=C)
    
        # Use the cross-validation accuracy to compare the models' performance
        accuracy =  ? (clf, X, y).mean()

    Sources4521

    3.Parallelizing single-node tuning with SparkTrials

    SparkTrials distributes a Hyperopt run without any other change to your code. The driver generates trials, and each trial, meaning one model fitted on one setting, runs as a single-task Spark job on a worker. To use it, add one argument to fmin(). Wrapping the call in mlflow.start_run() turns on the automated MLflow tracking:

    Distributing fmin across Spark workers with MLflow trackingpython
    spark_trials = SparkTrials()
    
    with mlflow.start_run():
      argmin = fmin(
        fn=objective,
        space=search_space,
        algo=algo,
        max_evals=16,
        trials=spark_trials)
    Choosing the trials argument for fmin()
    AspectSparkTrialsTrials
    Use withSingle-machine algorithms such as scikit-learnDistributed algorithms such as MLlib or Horovod
    Where trials are evaluatedOn Spark worker nodesOn the cluster driver, so the algorithm can distribute itself
    MLflow loggingAutomated tracking when fmin() runs inside mlflow.start_run()No automatic logging; call MLflow manually

    SparkTrials takes two optional arguments. parallelism sets how many trials run concurrently; it defaults to the number of available executors and is capped at 128. timeout sets the maximum number of seconds fmin() may run. Parallelism has a cost. TPE proposes new trials from past results, so for a fixed max_evals, higher parallelism finishes sooner but gives each proposal fewer completed results to learn from. Don't use SparkTrials on autoscaling clusters: parallelism is fixed when execution starts, so nodes added later go unused. And for very short trials, Spark overhead can cancel out any speedup.

    Checkpoint 5 of 6· Exam question

    A team wants `GridSearchCV` to search over both `n_components` for a `PCA` step and `C` for an `SVC` step inside the same `Pipeline`, where those steps are named `reduce_dim` and `clf` respectively. Which `param_grid` correctly targets these nested parameters?

    Checkpoint 6 of 6· Check yourself

    With max_evals fixed at 64, a teammate raises SparkTrials parallelism from 8 to 64 to "get the best result". What is the likely trade-off?

    Sources465

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.SparkTrials is the right way to speed up Hyperopt for any model, including MLlib.Why is that wrong?

      SparkTrials is for single-machine libraries such as scikit-learn. With distributed algorithms like MLlib or Horovod, use the default Trials class.

      Covered in Parallelizing single-node tuning with SparkTrials

    2. 2.The objective function should return accuracy so Hyperopt can find the highest value.Why is that wrong?

      fmin minimizes the loss, so a higher-is-better score must be returned as a negative value.

      Covered in Tuning with Hyperopt fmin

    3. 3.Setting parallelism as high as possible always gives the best tuning result.Why is that wrong?

      For a fixed max_evals, higher parallelism is faster, but lower parallelism lets TPE learn from more completed trials.

      Covered in Parallelizing single-node tuning with SparkTrials

    Practise it for real

    Tune an SVM's C parameter with Hyperopt on a Databricks Runtime ML cluster at 16.4 LTS ML or below, and track every trial in MLflow.

    1. 1.Write an objective(C) that builds SVC(C=C), scores it with cross_val_score(clf, X, y).mean(), and returns {'loss': -accuracy, 'status': STATUS_OK}.

      Why: Hyperopt minimizes loss, so negating accuracy makes "better" mean "lower".

      You should see: Calling objective(1.0) returns a dictionary with a negative loss.

    2. 2.Define search_space = hp.lognormal('C', 0, 1.0) and set algo=tpe.suggest.

      Why: TPE adapts to past results, unlike rand.suggest.

      You should see: No output; the space and algorithm are ready to pass to fmin.

    3. 3.Run fmin(fn=objective, space=search_space, algo=algo, max_evals=16) without a trials argument.

      Why: This gives a single-machine baseline before you distribute anything.

      You should see: A printed dictionary with the best value found for C.

    4. 4.Create spark_trials = SparkTrials() and rerun fmin inside with mlflow.start_run(): passing trials=spark_trials.

      Why: SparkTrials sends trials to the workers, and the run context turns on automated MLflow tracking.

      You should see: The notebook experiment shows the runs. For an objective this small, Spark overhead may make it slower than the baseline.

    Stuck? Get a nudge

    If the distributed run is not faster, look at trial duration: Spark job overhead dominates when each trial lasts only a few seconds.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Use the cross-validation accuracy to compare the models' performance”
      ↩︎ Cross-validation versus a single validation split
      “hyperopt.rand.suggest: Random search, a non-adaptive approach that samples over the search space”
      ↩︎ Tuning with Hyperopt fmin
      “Hyperopt tries to minimize the objective function.”
      ↩︎ Exam trap 2
      “Run the tuning algorithm with Hyperopt fmin()”
      ↩︎ Checkpoint
    2. 2.
      “With MLlib automated MLflow tracking, when you run tuning code that uses CrossValidator or TrainValidationSplit, hyperparameters and evaluation metrics are automatically logged in MLflow.”
      ↩︎ Cross-validation versus a single validation split
      “Hyperopt is not included in Databricks Runtime for Machine Learning after 16.4 LTS ML.”
      ↩︎ Tuning with Hyperopt fmin
    3. 3.
      “The performance measure reported by k-fold cross-validation is then the average of the values computed in the loop.”
      ↩︎ Cross-validation versus a single validation split
      “by splitting the data, fitting a model and computing the score 5 consecutive times (with different splits each time)”
      ↩︎ Prediction
      “This approach can be computationally expensive, but does not waste too much data (as is the case when fixing an arbitrary validation set)”
      ↩︎ Checkpoint
    4. 4.
      “Number of hyperparameter settings to try (the number of models to fit).”
      ↩︎ Tuning with Hyperopt fmin
      “SparkTrials accelerates single-machine tuning by distributing trials to Spark workers.”
      ↩︎ Parallelizing single-node tuning with SparkTrials
      “Use Trials when you call distributed training algorithms such as MLlib methods or Horovod in the objective function.”
      ↩︎ Exam trap 1
      “lower parallelism may lead to better results since each iteration has access to more past results.”
      ↩︎ Exam trap 3
      “For models created with distributed ML algorithms such as MLlib or Horovod, do not use SparkTrials.”
      ↩︎ Prediction
      “lower parallelism may lead to better results since each iteration has access to more past results.”
      ↩︎ Checkpoint
    5. 5.
      “Bayesian approaches can be much more efficient than grid search and random search.”
      ↩︎ Tuning with Hyperopt fmin
      “When you use hp.choice(), Hyperopt returns the index of the choice list.”
      ↩︎ Tuning with Hyperopt fmin
      “Do not use SparkTrials on autoscaling clusters.”
      ↩︎ Parallelizing single-node tuning with SparkTrials

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.