What you will be able to do
- Explain how adding model complexity lowers bias and raises variance, and why validation error follows a U-shaped curve
- Diagnose underfitting and overfitting by comparing training scores with validation or cross-validation scores
- Name the hyperparameters that control complexity (tree depth, number of estimators, regularization strength, early stopping) and say which way to turn each one
Key concept
Bias-variance tradeoff — As a model gets more flexible, its error from wrong simplifying assumptions (bias) goes down, but its sensitivity to the particular training sample (variance) goes up. The model that performs best on unseen data sits between too simple and too complex. It is rarely the one with the best training score.
1.What model complexity means and why more is not always better
Model complexity is how flexible a model is: how many different patterns it can represent. A linear regression with a few features is low in complexity, because it can only draw a flat surface through the data. A decision tree allowed to grow 32 levels deep is high in complexity, because it can carve the feature space into thousands of small regions and give each region its own prediction. Gradient-boosted ensembles get more complex as you add trees and let each tree grow deeper.
Complexity is usually not fixed by the algorithm. You set it with hyperparameters. The Databricks Optuna example tunes it directly: it searches a random forest's max_depth between 2 and 32, and it searches an SVR's C, which sets how strongly the model is regularized. Both searches are really asking the same question: how complex should this model be?
regressor_name = trial.suggest_categorical('classifier', ['SVR', 'RandomForest'])
if regressor_name == 'SVR':
svr_c = trial.suggest_float('svr_c', 1e-10, 1e10, log=True)
regressor_obj = sklearn.svm.SVR(C=svr_c)
else:
rf_max_depth = trial.suggest_int('rf_max_depth', 2, 32)
regressor_obj = sklearn.ensemble.RandomForestRegressor(max_depth=rf_max_depth)Why not just pick the most flexible setting? Because a flexible enough model can memorize its training data. The scikit-learn guide puts it this way: a model that only repeats the labels it has already seen gets a perfect training score but predicts nothing useful on new data. That is overfitting. What matters is performance on data the model has not seen, and past a certain point more complexity makes that worse.
Checkpoint 1 of 7· Check yourself
A deep decision tree scores 100% accuracy on its training data. What can you conclude about how it will perform in production?
A model that only memorized its training labels would also score perfectly on that data. Only held-out data can tell you whether it generalizes, so the training score alone settles nothing.
“would have a perfect score but would fail to predict anything useful on yet-unseen data”Source: scikit-learn.org
2.The bias-variance tradeoff
The error a model makes on unseen data can be split into three parts:
- Bias: error from assumptions that are too simple. A straight line fitted to a curved relationship misses the curve every time, however much data you give it. High-bias models underfit: they do poorly on the training data and on validation data. - Variance: error from being too sensitive to the particular training sample. Retrain a very deep tree on a slightly different sample and its predictions change a lot, because it has fitted the noise. High-variance models overfit: they do very well on training data and noticeably worse on validation data. - Irreducible error: noise in the target itself. No model can remove it.
Complexity moves bias and variance in opposite directions. As complexity goes up, bias falls and variance rises. Training error keeps going down. Validation error first falls, while the drop in bias outweighs the rise in variance, and then climbs again once variance takes over. The result is the familiar U-shaped validation curve, and the best model sits at the bottom of the U.
| Hyperparameter | Turning it up tends to | Symptom if pushed too far |
|---|---|---|
| max_depth | Lower bias: each tree can fit finer interactions | High variance: training score far above validation score |
| n_estimators | Lower bias in boosting: each extra tree corrects more of the remaining training error | Validation metric levels off and then gets worse while training keeps improving |
| learning_rate | Faster fitting: each boosting step makes a bigger correction | Overshooting and overfitting, often needing fewer trees or stronger regularization |
It reduced bias, which is why training F1 improved. It also added variance, and the extra variance cancelled most of the gain on unseen data. That is why validation F1 hardly changed. The next step is to rein in complexity (shallower trees, fewer estimators, more regularization), not to add more.
Checkpoint 2 of 7· Exam question
A data scientist trains a decision tree classifier on a churn dataset without setting `max_depth`, achieving 99% training accuracy but only 61% accuracy on a held-out validation set. Which explanation and next step best addresses this gap?
Correct answer: A — The tree has grown deep enough to memorize noise in the training data, producing high variance; limiting `max_depth` or increasing `min_samples_leaf` will reduce complexity and close the gap.
- A. An unconstrained tree can keep splitting until each leaf fits individual training points, which drives training accuracy toward 100% while variance on unseen data grows. Reducing `max_depth` or raising `min_samples_leaf` caps this complexity and typically narrows the accuracy gap.
- B. A 99% training accuracy rules out underfitting; a high-bias model would perform poorly on both the training and validation sets, not just the validation set. Adding more feature interactions would only make an already-overcomplex tree fit training noise even more closely.
- C. A distribution mismatch between splits is a data problem, not a complexity problem, and nothing in the scenario suggests the splits were drawn unevenly. The described pattern — near-perfect training accuracy with much weaker validation accuracy — is the textbook signature of excess model complexity, not a split issue.
- D. Enlarging the validation set changes how precisely the gap is measured, but it does not change the tree's tendency to memorize training data. The underlying cause is the model's unconstrained complexity, which must be reduced directly.
Sources2
3.Diagnosing underfitting and overfitting
You diagnose bias and variance by comparing two numbers: the score on the training data and the score on data the model has not seen.
- Both scores poor, small gap: high bias, so the model is underfitting. Add complexity, for example a more flexible algorithm, deeper trees, more features, or less regularization. - Training score excellent, validation score much worse: high variance, so the model is overfitting. Reduce complexity, regularize harder, stop training earlier, or get more training data. - Both scores good, small gap: you are near the bottom of the U.
The unseen data has to be kept separate. If you keep adjusting hyperparameters until the test score looks good, the test set stops being unseen. Knowledge about it leaks into your choices, and the score no longer measures generalization. That is why a separate validation set, or k-fold cross-validation on the training data, is used to choose complexity. The test set is kept for one final check.
In scikit-learn, cross_validate can return training scores next to validation scores, but only when you ask for them, because computing them costs extra time.
Checkpoint 3 of 7· Check yourself
You call scikit-learn's cross_validate to check for overfitting, but the result has only test scores. What do you need to change?
cross_validate leaves out training scores by default to save time. Setting return_train_score=True adds them, so you can compare training and validation performance fold by fold.
“return_train_score is set to False by default to save computation time.”Source: scikit-learn.org
Checkpoint 4 of 7· Exam question
A team fits an ordinary linear regression model to predict house prices from a dataset with clearly nonlinear relationships between features and price. Both training R-squared (0.34) and validation R-squared (0.33) are low and close to each other. What does this pattern indicate, and what should the team try next?
Correct answer: A — The model has high bias because it is too simple to capture the nonlinear relationships; increasing complexity by adding polynomial features or switching to a tree-based model should improve both scores.
- A. Low training and validation scores that sit close together indicate the model is too simple to represent the underlying nonlinear pattern, which is the definition of high bias. Raising complexity through polynomial features or a model class that captures nonlinearity, such as a tree-based method, addresses the true source of the poor fit.
- B. High variance shows up as a large gap between a high training score and a much lower validation score, not as two scores that are both low and close together. Reducing complexity further would only make an already underfit model fit the data even worse.
- C. Outliers can degrade R-squared, but they do not typically produce training and validation scores that are both uniformly low and nearly identical, which instead points to a systematic capacity limitation. Abandoning regression outright ignores that the real fix is matching the model's complexity to the data's nonlinearity.
- D. A flawed split would usually produce a mismatch between training and validation performance rather than two consistently low scores that agree with each other. The consistency between the scores points to the model's limited capacity, not to how the data was partitioned.
4.Controlling complexity in practice
Once you know which side of the U you are on, you have a few practical controls.
Regularization. Regularization adds a penalty for large coefficients, which pulls the model toward simpler solutions. SparkR's glm exposes lambda as the regularization parameter and alpha as the elastic-net mixing parameter, which blends L1 (lasso) and L2 (ridge) penalties. A larger lambda means a simpler model: more bias, less variance. In scikit-learn's SVR and SVC the parameter C works the other way round: a smaller C means stronger regularization.
Structural limits. For tree models, cap max_depth, limit how many boosting trees you add, and lower the learning rate. The Databricks getting-started tutorial tunes exactly these three for a GradientBoostingClassifier, and the narrow max_depth range of 2 to 5 is itself a deliberate cap on complexity.
params = {
'n_estimators': trial.suggest_int('n_estimators', 20, 1000),
'learning_rate': trial.suggest_float('learning_rate', 0.05, 1.0, log=True),
'max_depth': trial.suggest_int('max_depth', 2, 5),
}Early stopping. For models trained step by step, such as neural networks and boosting, complexity grows with every epoch or round. Early stopping watches a validation metric and stops once it no longer improves, which in effect finds the bottom of the U automatically. Databricks recommends this over guessing a fixed number of epochs, and AutoML applies the same idea when it stops tuning once the validation metric levels off.
Let tuning choose, but score on validation data. Hyperparameter search with Optuna, grid search or random search is really a search along the complexity axis. It is only as honest as the score it optimizes, so that score has to come from validation folds and not from the training data.
Checkpoint 5 of 7· Fill the gap
Which constructor argument links the tuned depth value to the random forest's complexity?
rf_max_depth = trial.suggest_int('rf_max_depth', 2, 32)
regressor_obj = sklearn.ensemble.RandomForestRegressor( ? =rf_max_depth)max_depth caps how many levels each tree can grow, so it is a direct control on complexity. A bigger value means lower bias and higher variance.
Source: docs.databricks.comCheckpoint 6 of 7· Match them up
Match each complexity control to what it does
Tap a term, then the definition that fits it.
lambda and alpha are glm's regularization controls, max_depth limits tree structure, and early stopping limits how long iterative training runs. Each one reduces variance at the cost of some bias.
“lambda: Numeric, Regularization parameter”Source: docs.databricks.com
Checkpoint 7 of 7· Exam question
A data scientist sets `n_neighbors=1` for a k-nearest neighbors classifier and observes that predictions change drastically when a single training point near the decision boundary is removed. Which statement best explains this behavior and how to address it?
Correct answer: A — With `n_neighbors=1`, model complexity is effectively maximized because each prediction depends on one nearby point, producing high variance; increasing `n_neighbors` trades some variance for bias.
- A. A single-neighbor classifier draws a highly irregular decision boundary that hugs each individual training point, so removing or adding one point can flip nearby predictions — the hallmark of high variance. Increasing `n_neighbors` averages over more points, smoothing the boundary and accepting some bias in exchange for lower variance.
- B. With only one neighbor there is no averaging happening at all, so this describes the opposite regime from what `n_neighbors=1` actually produces. Decreasing `n_neighbors` below 1 is not possible and would not smooth anything even if it were.
- C. Feature scaling affects which points are considered nearest, but it does not remove the fundamental instability caused by basing every prediction on a single neighbor. Even on perfectly scaled features, a one-neighbor model remains maximally sensitive to individual training points.
- D. k-nearest neighbors is a classic example of a non-parametric method, and `n_neighbors` is precisely the hyperparameter that controls its effective complexity. The observed instability is expected behavior for this setting, not evidence of a pipeline defect.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The model with the highest training score is the best model.Why is that wrong?
A complex enough model can memorize its training data and still fail on new data. Compare candidates on held-out validation or cross-validation scores.
Covered in What model complexity means and why more is not always better
2.It is fine to adjust complexity hyperparameters until the test-set score is as high as possible.Why is that wrong?
Repeatedly tuning against the test set leaks information from it into your choices, so its score stops measuring generalization. Tune on a validation set or with cross-validation, and use the test set once at the end.
Covered in Diagnosing underfitting and overfitting
3.Training for more epochs or boosting rounds always improves the model, because training loss keeps falling.Why is that wrong?
Training loss keeps falling while validation performance eventually gets worse as variance grows. Early stopping uses the validation metric to stop at the right point.
Covered in Controlling complexity in practice
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“rf_max_depth = trial.suggest_int('rf_max_depth', 2, 32)”
↩︎ What model complexity means and why more is not always better - 2.https://scikit-learn.org/stable/modules/cross_validation.htmlSecondary source
“This situation is called overfitting.”
↩︎ What model complexity means and why more is not always better“a model that would just repeat the labels of the samples that it has just seen would have a perfect score”
↩︎ The bias-variance tradeoff“yet another part of the dataset can be held out as a so-called “validation set””
↩︎ Diagnosing underfitting and overfitting“a model that would just repeat the labels of the samples that it has just seen would have a perfect score”
↩︎ Key concept“would have a perfect score but would fail to predict anything useful on yet-unseen data”
↩︎ Exam trap 1“knowledge about the test set can “leak” into the model and evaluation metrics no longer report on generalization performance.”
↩︎ Exam trap 2“would have a perfect score but would fail to predict anything useful on yet-unseen data”
↩︎ Checkpoint“This situation is called overfitting.”
↩︎ Prediction“return_train_score is set to False by default to save computation time.”
↩︎ Checkpoint - 3.
“AutoML incorporates early stopping; it stops training and tuning models if the validation metric is no longer improving.”
↩︎ Diagnosing underfitting and overfitting - 4.
“alpha: Numeric, Elastic-net mixing parameter”
↩︎ Controlling complexity in practice“lambda: Numeric, Regularization parameter”
↩︎ Checkpoint - 5.
“This is a better approach than guessing at a good number of epochs to complete.”
↩︎ Controlling complexity in practice“Early stopping monitors the value of a metric calculated on the validation set and stops training when the metric stops improving.”
↩︎ Exam trap 3