CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 42/48

    Bias-Variance Tradeoff and Model Complexity

    Assess the impact of model complexity and the bias variance tradeoff on model performance

    13 min read
    2.08% of exam
    5 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain how adding model complexity lowers bias and raises variance, and why validation error follows a U-shaped curve
    • Diagnose underfitting and overfitting by comparing training scores with validation or cross-validation scores
    • Name the hyperparameters that control complexity (tree depth, number of estimators, regularization strength, early stopping) and say which way to turn each one

    Key concept

    Bias-variance tradeoff — As a model gets more flexible, its error from wrong simplifying assumptions (bias) goes down, but its sensitivity to the particular training sample (variance) goes up. The model that performs best on unseen data sits between too simple and too complex. It is rarely the one with the best training score.

    1.What model complexity means and why more is not always better

    Model complexity is how flexible a model is: how many different patterns it can represent. A linear regression with a few features is low in complexity, because it can only draw a flat surface through the data. A decision tree allowed to grow 32 levels deep is high in complexity, because it can carve the feature space into thousands of small regions and give each region its own prediction. Gradient-boosted ensembles get more complex as you add trees and let each tree grow deeper.

    Complexity is usually not fixed by the algorithm. You set it with hyperparameters. The Databricks Optuna example tunes it directly: it searches a random forest's max_depth between 2 and 32, and it searches an SVR's C, which sets how strongly the model is regularized. Both searches are really asking the same question: how complex should this model be?

    A Databricks Optuna objective that puts complexity hyperparameters (SVR's C and the random forest's max_depth) into the search spacepython
        regressor_name = trial.suggest_categorical('classifier', ['SVR', 'RandomForest'])
        if regressor_name == 'SVR':
            svr_c = trial.suggest_float('svr_c', 1e-10, 1e10, log=True)
            regressor_obj = sklearn.svm.SVR(C=svr_c)
        else:
            rf_max_depth = trial.suggest_int('rf_max_depth', 2, 32)
            regressor_obj = sklearn.ensemble.RandomForestRegressor(max_depth=rf_max_depth)

    Why not just pick the most flexible setting? Because a flexible enough model can memorize its training data. The scikit-learn guide puts it this way: a model that only repeats the labels it has already seen gets a perfect training score but predicts nothing useful on new data. That is overfitting. What matters is performance on data the model has not seen, and past a certain point more complexity makes that worse.

    Checkpoint 1 of 7· Check yourself

    A deep decision tree scores 100% accuracy on its training data. What can you conclude about how it will perform in production?

    Sources12

    2.The bias-variance tradeoff

    The error a model makes on unseen data can be split into three parts:

    - Bias: error from assumptions that are too simple. A straight line fitted to a curved relationship misses the curve every time, however much data you give it. High-bias models underfit: they do poorly on the training data and on validation data. - Variance: error from being too sensitive to the particular training sample. Retrain a very deep tree on a slightly different sample and its predictions change a lot, because it has fitted the noise. High-variance models overfit: they do very well on training data and noticeably worse on validation data. - Irreducible error: noise in the target itself. No model can remove it.

    Complexity moves bias and variance in opposite directions. As complexity goes up, bias falls and variance rises. Training error keeps going down. Validation error first falls, while the drop in bias outweighs the rise in variance, and then climbs again once variance takes over. The result is the familiar U-shaped validation curve, and the best model sits at the bottom of the U.

    How the three main tree-ensemble complexity knobs from the Databricks tutorial move bias and variance
    HyperparameterTurning it up tends toSymptom if pushed too far
    max_depthLower bias: each tree can fit finer interactionsHigh variance: training score far above validation score
    n_estimatorsLower bias in boosting: each extra tree corrects more of the remaining training errorValidation metric levels off and then gets worse while training keeps improving
    learning_rateFaster fitting: each boosting step makes a bigger correctionOvershooting and overfitting, often needing fewer trees or stronger regularization

    Checkpoint 2 of 7· Exam question

    A data scientist trains a decision tree classifier on a churn dataset without setting `max_depth`, achieving 99% training accuracy but only 61% accuracy on a held-out validation set. Which explanation and next step best addresses this gap?

    Sources2

    3.Diagnosing underfitting and overfitting

    You diagnose bias and variance by comparing two numbers: the score on the training data and the score on data the model has not seen.

    - Both scores poor, small gap: high bias, so the model is underfitting. Add complexity, for example a more flexible algorithm, deeper trees, more features, or less regularization. - Training score excellent, validation score much worse: high variance, so the model is overfitting. Reduce complexity, regularize harder, stop training earlier, or get more training data. - Both scores good, small gap: you are near the bottom of the U.

    The unseen data has to be kept separate. If you keep adjusting hyperparameters until the test score looks good, the test set stops being unseen. Knowledge about it leaks into your choices, and the score no longer measures generalization. That is why a separate validation set, or k-fold cross-validation on the training data, is used to choose complexity. The test set is kept for one final check.

    In scikit-learn, cross_validate can return training scores next to validation scores, but only when you ask for them, because computing them costs extra time.

    Checkpoint 3 of 7· Check yourself

    You call scikit-learn's cross_validate to check for overfitting, but the result has only test scores. What do you need to change?

    Checkpoint 4 of 7· Exam question

    A team fits an ordinary linear regression model to predict house prices from a dataset with clearly nonlinear relationships between features and price. Both training R-squared (0.34) and validation R-squared (0.33) are low and close to each other. What does this pattern indicate, and what should the team try next?

    Sources32

    4.Controlling complexity in practice

    Once you know which side of the U you are on, you have a few practical controls.

    Regularization. Regularization adds a penalty for large coefficients, which pulls the model toward simpler solutions. SparkR's glm exposes lambda as the regularization parameter and alpha as the elastic-net mixing parameter, which blends L1 (lasso) and L2 (ridge) penalties. A larger lambda means a simpler model: more bias, less variance. In scikit-learn's SVR and SVC the parameter C works the other way round: a smaller C means stronger regularization.

    Structural limits. For tree models, cap max_depth, limit how many boosting trees you add, and lower the learning rate. The Databricks getting-started tutorial tunes exactly these three for a GradientBoostingClassifier, and the narrow max_depth range of 2 to 5 is itself a deliberate cap on complexity.

    Search space from the Databricks getting-started tutorial: n_estimators, learning_rate and max_depth together set the gradient-boosted model's complexitypython
        params = {
          'n_estimators': trial.suggest_int('n_estimators', 20, 1000),
          'learning_rate': trial.suggest_float('learning_rate', 0.05, 1.0, log=True),
          'max_depth': trial.suggest_int('max_depth', 2, 5),
        }

    Early stopping. For models trained step by step, such as neural networks and boosting, complexity grows with every epoch or round. Early stopping watches a validation metric and stops once it no longer improves, which in effect finds the bottom of the U automatically. Databricks recommends this over guessing a fixed number of epochs, and AutoML applies the same idea when it stops tuning once the validation metric levels off.

    Let tuning choose, but score on validation data. Hyperparameter search with Optuna, grid search or random search is really a search along the complexity axis. It is only as honest as the score it optimizes, so that score has to come from validation folds and not from the training data.

    Checkpoint 5 of 7· Fill the gap

    Which constructor argument links the tuned depth value to the random forest's complexity?

    rf_max_depth = trial.suggest_int('rf_max_depth', 2, 32)
            regressor_obj = sklearn.ensemble.RandomForestRegressor( ? =rf_max_depth)

    Checkpoint 6 of 7· Match them up

    Match each complexity control to what it does

    Tap a term, then the definition that fits it.

    Checkpoint 7 of 7· Exam question

    A data scientist sets `n_neighbors=1` for a k-nearest neighbors classifier and observes that predictions change drastically when a single training point near the decision boundary is removed. Which statement best explains this behavior and how to address it?

    Sources45

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The model with the highest training score is the best model.Why is that wrong?

      A complex enough model can memorize its training data and still fail on new data. Compare candidates on held-out validation or cross-validation scores.

      Covered in What model complexity means and why more is not always better

    2. 2.It is fine to adjust complexity hyperparameters until the test-set score is as high as possible.Why is that wrong?

      Repeatedly tuning against the test set leaks information from it into your choices, so its score stops measuring generalization. Tune on a validation set or with cross-validation, and use the test set once at the end.

      Covered in Diagnosing underfitting and overfitting

    3. 3.Training for more epochs or boosting rounds always improves the model, because training loss keeps falling.Why is that wrong?

      Training loss keeps falling while validation performance eventually gets worse as variance grows. Early stopping uses the validation metric to stop at the right point.

      Covered in Controlling complexity in practice

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 2.
      “a model that would just repeat the labels of the samples that it has just seen would have a perfect score”
      ↩︎ The bias-variance tradeoff
      “yet another part of the dataset can be held out as a so-called “validation set””
      ↩︎ Diagnosing underfitting and overfitting
      “a model that would just repeat the labels of the samples that it has just seen would have a perfect score”
      ↩︎ Key concept
      “would have a perfect score but would fail to predict anything useful on yet-unseen data”
      ↩︎ Exam trap 1
      “knowledge about the test set can “leak” into the model and evaluation metrics no longer report on generalization performance.”
      ↩︎ Exam trap 2
      “would have a perfect score but would fail to predict anything useful on yet-unseen data”
      ↩︎ Checkpoint
      “This situation is called overfitting.”
      ↩︎ Prediction
      “return_train_score is set to False by default to save computation time.”
      ↩︎ Checkpoint
    2. 3.
      “AutoML incorporates early stopping; it stops training and tuning models if the validation metric is no longer improving.”
      ↩︎ Diagnosing underfitting and overfitting
    3. 4.
      “alpha: Numeric, Elastic-net mixing parameter”
      ↩︎ Controlling complexity in practice
      “lambda: Numeric, Regularization parameter”
      ↩︎ Checkpoint
    4. 5.
      “This is a better approach than guessing at a good number of epochs to complete.”
      ↩︎ Controlling complexity in practice
      “Early stopping monitors the value of a metric calculated on the validation set and stops training when the metric stops improving.”
      ↩︎ Exam trap 3

    Spotted a mistake, or was something unclear? Tell us.