CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 28/48

    Algorithm Selection: Linear, Logistic and Tree-Based Models

    Use ML foundations to select the appropriate algorithm for a given model scenario

    12 min read
    2.08% of exam
    8 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Work out from the prediction target whether a scenario is a classification, regression or forecasting problem
    • Choose between linear regression and logistic regression based on the type of the target
    • Explain when tree-based models such as random forests and gradient boosted trees are a better starting point than linear models
    • Use the bias-variance tradeoff and stakeholder requirements such as explainability to justify an algorithm choice

    Key concept

    The prediction target sets the problem class — You can't pick an algorithm until you know what kind of value the model must output: a category, a continuous number, or a future value in a time series. That problem class narrows the candidate algorithms. Data shape, interpretability and cost then decide among them.

    1.Start from the target, not the algorithm

    Exam scenarios about algorithm selection usually hide the answer in one detail: what the model has to predict. Databricks' ML lifecycle guidance makes this the first scoping question, before any data preparation or training. Ask what the target column holds. Then ask which class of ML problem that implies.

    Databricks frames three supervised problem classes for tabular data. Each one maps to a different kind of target. Classification predicts the label or category of an input: churn or no churn, fraud or not fraud, one of five product types. Regression predicts a continuous numeric value, such as price, revenue or temperature. Forecasting predicts values from time-series data. In forecasting, the order of the observations and a time column are part of the problem itself, not just another feature. Settle the class first. An algorithm that is excellent for one class is often simply the wrong tool for another.

    Checkpoint 1 of 7· Check yourself

    A model must decide, for each incoming insurance claim, whether to route it to manual review. Which problem class is this?

    Checkpoint 2 of 7· Exam question

    A data scientist on Databricks is predicting a continuous target, total monthly revenue, from five numeric features. Diagnostic plots show a roughly straight-line relationship between each feature and revenue, and the business asks for a model whose coefficients can be quoted directly to stakeholders as "revenue changes by X per unit of feature Y." Which model best fits this scenario?

    Sources123

    2.Linear and logistic regression: the linear family

    Linear regression and logistic regression are closely related. Both learn one weight per feature and combine the features additively. The difference is what they predict. Linear regression outputs a continuous number, so it's a regression algorithm. Logistic regression passes the same weighted sum through a logistic function to get a class probability. Despite its name, it's a classification algorithm. The Databricks SparkR tutorial makes the relationship explicit: both are fitted with the same glm function, and only the family argument changes.

    Linear regression in the Databricks glm tutorial: a continuous target (price) with family set to gaussianr
    # Family = "gaussian" to train a linear regression model
    lrModel <- glm(price ~ ., data = trainingData, family = "gaussian")

    The same tutorial then predicts a diamond's cut, which is a category, using only two cut values. MLlib's logistic regression in that example supports binary classification. When should you reach for the linear family? Choose it when the relationship between features and target is roughly linear (or can be made linear with feature engineering), when you need coefficients that are easy to explain, or when you want a fast, low-variance baseline. Its main weakness is high bias: it can't capture interactions or thresholds unless you engineer them in as features. Linear models are also sensitive to feature scale, so they usually need standardization first.

    Checkpoint 3 of 7· Fill the gap

    The tutorial trains a model to predict a diamond's cut (one of two categories). Which family value completes the call?

    logrModel <- glm(cut ~ price + color + clarity + depth, data = trainingDataSub, family = " ? ")

    Sources45

    3.Tree-based models and the bias-variance tradeoff

    Tree-based models split the feature space with a series of if/then thresholds. Because of this, they capture non-linear relationships and feature interactions without manual feature engineering. They also don't depend on feature scale the way linear models do. A single decision tree is easy to read, but it overfits easily: a deep tree has low bias and high variance. Ensembles fix this. Random forests average many decorrelated trees to cut variance. Gradient boosted trees, such as XGBoost and LightGBM, add shallow trees one after another, each correcting the errors of the trees before it. Both work for classification and for regression, which is why AutoML offers them for both problem types.

    Algorithms that Databricks AutoML trains and evaluates, by problem class
    AlgorithmClassificationRegressionForecasting
    Decision treesYesYesNo
    Random forestsYesYesNo
    Logistic regressionYesNoNo
    Linear regression with stochastic gradient descentNoYesNo
    XGBoost / LightGBMYesYesNo
    Prophet / Auto-ARIMANoNoYes

    That table also reflects the bias-variance tradeoff. Linear models sit at the high-bias, low-variance end: they're stable but may underfit. Single deep trees sit at the low-bias, high-variance end: they're flexible but may overfit. Ensembles aim for the middle. When an exam scenario says a model does well on training data but poorly on new data, it is describing high variance. The fix is a less complex model or one that reduces variance, such as a shallower tree or more averaging. When a model does poorly even on training data, it is describing high bias. The fix is a more flexible algorithm or richer features.

    Checkpoint 4 of 7· Check yourself

    Which statement about how Databricks AutoML uses decision trees and random forests is correct?

    Checkpoint 5 of 7· Exam question

    A team is building a churn model where the target is binary (churned vs. retained) and the deployment requirement is a well-calibrated probability of churn for each customer, since the marketing team ranks customers by risk score to prioritize outreach. Which algorithm choice best satisfies this requirement out of the box?

    Sources16

    4.Requirements narrow the choice; they don't dictate it

    Once the problem class is fixed, practical requirements break the tie. The lifecycle guidance lists several scoping questions: the success metrics, the serving requirements for latency and throughput, and which stakeholders must sign off and what they need in terms of explainability. A regulated lending decision may favor logistic regression, because its coefficients can be explained directly. A recommendation feature with a tight latency budget may favor a compact model over a large ensemble. Still, the guidance is explicit that these requirements don't determine one specific ML method. They narrow the field, and then you compare candidates empirically.

    Checkpoint 6 of 7· Check yourself

    A team has defined its prediction target, success metric and latency requirements. According to Databricks' lifecycle guidance, what follows for the choice of ML method?

    Because several algorithms usually remain plausible, Databricks AutoML compares them for you. It trains and tunes candidates across several algorithm families and reports the best one for your chosen metric. It also generates a source notebook for each trial, so you can inspect what was tried. Even with AutoML, the reasoning on this page still matters: you pick the problem type, and you need to judge whether the winning model meets your explainability and serving requirements.

    Checkpoint 7 of 7· Exam question

    A dataset for predicting loan default contains a mix of numeric and high-cardinality categorical features, several strong non-linear interactions between features, and some missing values in a few numeric columns. A data scientist wants a model family that can capture these interactions and handle the categorical and missing data with minimal preprocessing. Which model family is the most appropriate starting point?

    Sources17

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Logistic regression has "regression" in its name, so it's the right choice for predicting a continuous value such as price.Why is that wrong?

      Logistic regression is a classification algorithm that models class probabilities. Continuous targets call for linear regression, which uses the gaussian family in glm. Logistic regression uses the binomial family.

      Covered in Linear and logistic regression: the linear family

    2. 2.Predicting next month's value of a time series is just ordinary regression, so any regression algorithm will do.Why is that wrong?

      Forecasting on time-series data is its own problem class, with dedicated algorithms such as Prophet and Auto-ARIMA. The time ordering is part of the problem.

      Covered in Start from the target, not the algorithm

    3. 3.Once metrics and serving requirements are defined, they determine the one correct algorithm.Why is that wrong?

      Requirements narrow the candidates but don't dictate a method. You still compare approaches, often starting with something simpler like gradient boosted trees.

      Covered in Requirements narrow the choice; they don't dictate it

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “What is the prediction target, and what class of ML problem does that imply”
      ↩︎ Start from the target, not the algorithm
      “You might begin modeling with a simpler approach like gradient boosted trees, and later decide that more powerful deep learning methods are needed.”
      ↩︎ Tree-based models and the bias-variance tradeoff
      “Which stakeholders must sign off on production deployments? What are their requirements around explainability?”
      ↩︎ Requirements narrow the choice; they don't dictate it
      “What is the prediction target, and what class of ML problem does that imply”
      ↩︎ Key concept
      “The preceding requirements do not dictate the specific ML method.”
      ↩︎ Exam trap 3
      “The preceding requirements do not dictate the specific ML method.”
      ↩︎ Checkpoint
    2. 2.
      “Use AutoML to automatically find the best classification algorithm and hyperparameter configuration to predict the label or category of a given input.”
      ↩︎ Start from the target, not the algorithm
    3. 3.
      “Use AutoML to automatically find the best regression algorithm and hyperparameter configuration to predict continuous numeric values.”
      ↩︎ Start from the target, not the algorithm
    4. 4.
      “Logistic regression in MLlib supports binary classification.”
      ↩︎ Linear and logistic regression: the linear family
      “family: String, "gaussian" for linear regression or "binomial" for logistic regression”
      ↩︎ Exam trap 1
    5. 5.
      “In general, many learning algorithms such as linear models benefit from standardization of the data set”
      ↩︎ Linear and logistic regression: the linear family
    6. 6.
      “This notebook shows you how to use MLlib pipelines to perform a regression using gradient boosted trees to predict bike rental counts”
      ↩︎ Tree-based models and the bias-variance tradeoff
    7. 7.
      “Orchestrates distributed model training and hyperparameter tuning across multiple algorithms.”
      ↩︎ Requirements narrow the choice; they don't dictate it
      “For classification and regression models, the decision tree, random forests, logistic regression, and linear regression with stochastic gradient descent algorithms are based on scikit-learn.”
      ↩︎ Checkpoint

    Also cited

    Spotted a mistake, or was something unclear? Tell us.