What you will be able to do
- Work out from the prediction target whether a scenario is a classification, regression or forecasting problem
- Choose between linear regression and logistic regression based on the type of the target
- Explain when tree-based models such as random forests and gradient boosted trees are a better starting point than linear models
- Use the bias-variance tradeoff and stakeholder requirements such as explainability to justify an algorithm choice
Key concept
The prediction target sets the problem class — You can't pick an algorithm until you know what kind of value the model must output: a category, a continuous number, or a future value in a time series. That problem class narrows the candidate algorithms. Data shape, interpretability and cost then decide among them.
1.Start from the target, not the algorithm
Exam scenarios about algorithm selection usually hide the answer in one detail: what the model has to predict. Databricks' ML lifecycle guidance makes this the first scoping question, before any data preparation or training. Ask what the target column holds. Then ask which class of ML problem that implies.
Databricks frames three supervised problem classes for tabular data. Each one maps to a different kind of target. Classification predicts the label or category of an input: churn or no churn, fraud or not fraud, one of five product types. Regression predicts a continuous numeric value, such as price, revenue or temperature. Forecasting predicts values from time-series data. In forecasting, the order of the observations and a time column are part of the problem itself, not just another feature. Settle the class first. An algorithm that is excellent for one class is often simply the wrong tool for another.
Checkpoint 1 of 7· Check yourself
A model must decide, for each incoming insurance claim, whether to route it to manual review. Which problem class is this?
The decision is a category: route to review or don't. Some classifiers produce a probability internally, but the target is still a label, so this is a classification problem.
“Use AutoML to automatically find the best classification algorithm and hyperparameter configuration to predict the label or category of a given input.”Source: docs.databricks.com
Checkpoint 2 of 7· Exam question
A data scientist on Databricks is predicting a continuous target, total monthly revenue, from five numeric features. Diagnostic plots show a roughly straight-line relationship between each feature and revenue, and the business asks for a model whose coefficients can be quoted directly to stakeholders as "revenue changes by X per unit of feature Y." Which model best fits this scenario?
Correct answer: A — Linear regression, because it fits a continuous target with a linear relationship to the predictors and produces coefficients that directly quantify each feature's effect on the outcome.
- A. This is correct because linear regression is designed for continuous targets with linear predictor relationships, and its fitted coefficients have a direct, interpretable meaning: the expected change in the target per unit change in a feature.
- B. This is incorrect because logistic regression models the probability of a categorical (typically binary) outcome, not a continuous target like monthly revenue, so its coefficients describe log-odds rather than a direct revenue effect.
- C. This is incorrect because a random forest is an ensemble of decision trees without a single coefficient per feature; feature importance ranks relative influence but cannot be quoted as a per-unit revenue change the way a linear coefficient can.
- D. This is incorrect because k-means is an unsupervised clustering technique that groups observations by similarity and does not model or predict a target variable like revenue at all.
2.Linear and logistic regression: the linear family
Linear regression and logistic regression are closely related. Both learn one weight per feature and combine the features additively. The difference is what they predict. Linear regression outputs a continuous number, so it's a regression algorithm. Logistic regression passes the same weighted sum through a logistic function to get a class probability. Despite its name, it's a classification algorithm. The Databricks SparkR tutorial makes the relationship explicit: both are fitted with the same glm function, and only the family argument changes.
# Family = "gaussian" to train a linear regression model
lrModel <- glm(price ~ ., data = trainingData, family = "gaussian")The same tutorial then predicts a diamond's cut, which is a category, using only two cut values. MLlib's logistic regression in that example supports binary classification. When should you reach for the linear family? Choose it when the relationship between features and target is roughly linear (or can be made linear with feature engineering), when you need coefficients that are easy to explain, or when you want a fast, low-variance baseline. Its main weakness is high bias: it can't capture interactions or thresholds unless you engineer them in as features. Linear models are also sensitive to feature scale, so they usually need standardization first.
Checkpoint 3 of 7· Fill the gap
The tutorial trains a model to predict a diamond's cut (one of two categories). Which family value completes the call?
logrModel <- glm(cut ~ price + color + clarity + depth, data = trainingDataSub, family = " ? ")Cut is a two-class categorical target, so this is logistic regression, which in glm means family = "binomial". Gaussian would fit a linear regression for a continuous target.
Source: docs.databricks.com3.Tree-based models and the bias-variance tradeoff
Tree-based models split the feature space with a series of if/then thresholds. Because of this, they capture non-linear relationships and feature interactions without manual feature engineering. They also don't depend on feature scale the way linear models do. A single decision tree is easy to read, but it overfits easily: a deep tree has low bias and high variance. Ensembles fix this. Random forests average many decorrelated trees to cut variance. Gradient boosted trees, such as XGBoost and LightGBM, add shallow trees one after another, each correcting the errors of the trees before it. Both work for classification and for regression, which is why AutoML offers them for both problem types.
| Algorithm | Classification | Regression | Forecasting |
|---|---|---|---|
| Decision trees | Yes | Yes | No |
| Random forests | Yes | Yes | No |
| Logistic regression | Yes | No | No |
| Linear regression with stochastic gradient descent | No | Yes | No |
| XGBoost / LightGBM | Yes | Yes | No |
| Prophet / Auto-ARIMA | No | No | Yes |
That table also reflects the bias-variance tradeoff. Linear models sit at the high-bias, low-variance end: they're stable but may underfit. Single deep trees sit at the low-bias, high-variance end: they're flexible but may overfit. Ensembles aim for the middle. When an exam scenario says a model does well on training data but poorly on new data, it is describing high variance. The fix is a less complex model or one that reduces variance, such as a shallower tree or more averaging. When a model does poorly even on training data, it is describing high bias. The fix is a more flexible algorithm or richer features.
Checkpoint 4 of 7· Check yourself
Which statement about how Databricks AutoML uses decision trees and random forests is correct?
AutoML lists decision trees and random forests under both classification models and regression models, and those implementations come from scikit-learn.
“For classification and regression models, the decision tree, random forests, logistic regression, and linear regression with stochastic gradient descent algorithms are based on scikit-learn.”Source: docs.databricks.com
Checkpoint 5 of 7· Exam question
A team is building a churn model where the target is binary (churned vs. retained) and the deployment requirement is a well-calibrated probability of churn for each customer, since the marketing team ranks customers by risk score to prioritize outreach. Which algorithm choice best satisfies this requirement out of the box?
Correct answer: A — Logistic regression, because it models the log-odds of the positive class as a linear function of the features and its sigmoid output is a probability that tends to be well calibrated.
- A. This is correct because logistic regression is built specifically for binary targets, modeling log-odds linearly and passing them through a sigmoid to produce outputs that behave as class probabilities and are generally well calibrated.
- B. This is incorrect because linear regression can produce fitted values below 0 or above 1 for a 0/1 target, which are not valid probabilities and are not naturally calibrated for ranking risk.
- C. This is incorrect because a single unpruned decision tree tends to produce extreme leaf proportions (near 0 or 1) that overfit the training data and are poorly calibrated as probability estimates without additional smoothing.
- D. This is incorrect because hierarchical clustering is unsupervised and groups customers by feature similarity; it does not use the churn label and produces no probability output at all.
4.Requirements narrow the choice; they don't dictate it
Once the problem class is fixed, practical requirements break the tie. The lifecycle guidance lists several scoping questions: the success metrics, the serving requirements for latency and throughput, and which stakeholders must sign off and what they need in terms of explainability. A regulated lending decision may favor logistic regression, because its coefficients can be explained directly. A recommendation feature with a tight latency budget may favor a compact model over a large ensemble. Still, the guidance is explicit that these requirements don't determine one specific ML method. They narrow the field, and then you compare candidates empirically.
Checkpoint 6 of 7· Check yourself
A team has defined its prediction target, success metric and latency requirements. According to Databricks' lifecycle guidance, what follows for the choice of ML method?
Scoping narrows the options but doesn't pick the algorithm for you. You might start with gradient boosted trees and change approach later.
“The preceding requirements do not dictate the specific ML method.”Source: docs.databricks.com
Because several algorithms usually remain plausible, Databricks AutoML compares them for you. It trains and tunes candidates across several algorithm families and reports the best one for your chosen metric. It also generates a source notebook for each trial, so you can inspect what was tried. Even with AutoML, the reasoning on this page still matters: you pick the problem type, and you need to judge whether the winning model meets your explainability and serving requirements.
This is high variance, which is overfitting. Moving to a random forest, which averages many trees, or limiting the tree's depth reduces variance. Switching to a more complex model, such as a deeper tree, would make it worse.
Checkpoint 7 of 7· Exam question
A dataset for predicting loan default contains a mix of numeric and high-cardinality categorical features, several strong non-linear interactions between features, and some missing values in a few numeric columns. A data scientist wants a model family that can capture these interactions and handle the categorical and missing data with minimal preprocessing. Which model family is the most appropriate starting point?
Correct answer: A — Tree-based models such as gradient-boosted trees, because they split on raw feature values, naturally capture non-linear interactions, and tolerate categorical splits with little preprocessing.
- A. This is correct because tree-based models split on feature thresholds and categories directly, so they naturally represent non-linear interactions, handle high-cardinality categorical splits, and many gradient-boosting implementations handle missing values internally, all with minimal preprocessing.
- B. This is incorrect because manually engineering polynomial and interaction terms for many features does not scale well, still requires imputing missing values, and rarely captures the same complex, high-order interactions a tree ensemble finds automatically.
- C. This is incorrect because one-hot encoding a high-cardinality categorical feature creates a very sparse, wide feature space, and logistic regression's linear decision boundary still cannot represent strong non-linear interactions between features.
- D. This is incorrect because k-nearest neighbors is sensitive to feature scale and the curse of dimensionality with many categorical dummy columns, and it still requires missing values to be imputed before distances can be computed.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Logistic regression has "regression" in its name, so it's the right choice for predicting a continuous value such as price.Why is that wrong?
Logistic regression is a classification algorithm that models class probabilities. Continuous targets call for linear regression, which uses the gaussian family in glm. Logistic regression uses the binomial family.
Covered in Linear and logistic regression: the linear family
2.Predicting next month's value of a time series is just ordinary regression, so any regression algorithm will do.Why is that wrong?
Forecasting on time-series data is its own problem class, with dedicated algorithms such as Prophet and Auto-ARIMA. The time ordering is part of the problem.
Covered in Start from the target, not the algorithm
3.Once metrics and serving requirements are defined, they determine the one correct algorithm.Why is that wrong?
Requirements narrow the candidates but don't dictate a method. You still compare approaches, often starting with something simpler like gradient boosted trees.
Covered in Requirements narrow the choice; they don't dictate it
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“What is the prediction target, and what class of ML problem does that imply”
↩︎ Start from the target, not the algorithm“You might begin modeling with a simpler approach like gradient boosted trees, and later decide that more powerful deep learning methods are needed.”
↩︎ Tree-based models and the bias-variance tradeoff“Which stakeholders must sign off on production deployments? What are their requirements around explainability?”
↩︎ Requirements narrow the choice; they don't dictate it“What is the prediction target, and what class of ML problem does that imply”
↩︎ Key concept“The preceding requirements do not dictate the specific ML method.”
↩︎ Exam trap 3“The preceding requirements do not dictate the specific ML method.”
↩︎ Checkpoint - 2.
“Use AutoML to automatically find the best classification algorithm and hyperparameter configuration to predict the label or category of a given input.”
↩︎ Start from the target, not the algorithm - 3.
“Use AutoML to automatically find the best regression algorithm and hyperparameter configuration to predict continuous numeric values.”
↩︎ Start from the target, not the algorithm - 4.
“Logistic regression in MLlib supports binary classification.”
↩︎ Linear and logistic regression: the linear family“family: String, "gaussian" for linear regression or "binomial" for logistic regression”
↩︎ Exam trap 1 - 5.https://scikit-learn.org/stable/modules/preprocessing.htmlSecondary source
“In general, many learning algorithms such as linear models benefit from standardization of the data set”
↩︎ Linear and logistic regression: the linear family - 6.
“This notebook shows you how to use MLlib pipelines to perform a regression using gradient boosted trees to predict bike rental counts”
↩︎ Tree-based models and the bias-variance tradeoff - 7.
“Orchestrates distributed model training and hyperparameter tuning across multiple algorithms.”
↩︎ Requirements narrow the choice; they don't dictate it“For classification and regression models, the decision tree, random forests, logistic regression, and linear regression with stochastic gradient descent algorithms are based on scikit-learn.”
↩︎ Checkpoint
Also cited
“Use AutoML to automatically finding the best forecasting algorithm and hyperparameter configuration to predict values based on time-series data.”
↩︎ Exam trap 2