CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 39/48

    Regression error metrics: MSE, RMSE, MAE and their variants

    Use common regression metrics: RMSE, MAE, R-squared, etc.

    9 min read
    2.08% of exam
    5 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Compute MSE, RMSE and MAE by hand and with scikit-learn, and state the units each one is reported in
    • Explain why squared-error metrics target the mean of the response and absolute-error metrics target the median
    • Recognise when MAPE, MSLE, median absolute error or max_error answers a question that RMSE or MAE cannot
    • Name the regression metrics Databricks AutoML accepts as primary_metric

    Key concept

    Strictly consistent scoring function — Every regression metric quietly assumes which property of the outcome you are predicting. Squared-error metrics reward predicting the mean, and absolute-error metrics reward predicting the median. So you choose a metric by deciding first what the prediction is supposed to estimate.

    1.MSE and RMSE: averaging squared residuals

    Every regression metric starts from the residual: the true value y minus the prediction ŷ for each row. The metrics differ in how they turn a column of residuals into one number. The most common choice squares each residual and averages them. That gives the mean squared error (MSE), which scikit-learn describes as the expected value of the squared (quadratic) error. scikit-learn's own small example uses four rows. The residuals are 0.5, −0.5, 0 and −1, so the squared residuals are 0.25, 0.25, 0 and 1, and their average is 0.375.

    scikit-learn's mean_squared_error on a four-row examplepython
    >>> from sklearn.metrics import mean_squared_error
    >>> y_true = [3, -0.5, 2, 7]
    >>> y_pred = [2.5, 0.0, 2, 8]
    >>> mean_squared_error(y_true, y_pred)
    0.375

    Because MSE is in squared units, an MSE of 0.375 is hard to interpret directly. The root mean squared error (RMSE) is the square root of the MSE: √0.375 ≈ 0.61, in the same units as the target. scikit-learn provides it as root_mean_squared_error. You will also see it computed by hand. A Databricks TabFM tutorial wraps mean_squared_error in np.sqrt, and the SparkR glm tutorial builds RMSE directly from a column of errors. Both are the same formula: square the errors, average them, then take the square root.

    RMSE computed by hand in the Databricks SparkR glm tutorialr
    # Calculate RMSE
    head(select(errors, alias(sqrt(sum(errors$error^2 , na.rm = TRUE) / nrow(errors)), "RMSE")))

    Checkpoint 1 of 5· Fill the gap

    The Databricks TabFM tutorial computes RMSE from scikit-learn's mean_squared_error. Which NumPy function fills the blank?

    rmse = np. ? (mean_squared_error(y_test_reg, reg_pred))
    r2 = r2_score(y_test_reg, reg_pred)

    Sources123

    2.MAE, and what each metric is really measuring

    The mean absolute error (MAE) averages the absolute residuals instead of squaring them. On the same four rows, the absolute residuals are 0.5, 0.5, 0 and 1, so the MAE is 0.5. Like RMSE, MAE is in the units of the target. The two numbers differ because squaring weights large residuals more heavily. In the MSE, the single residual of 1 contributes 1 out of a total of 1.5. In the MAE, it contributes 1 out of a total of 2.

    scikit-learn's guide on choosing a scoring function answers this. The outcome Y is random for any given set of features, so a point prediction has to estimate some property of its distribution. Squared error is strictly consistent for the mean. Absolute error is strictly consistent for the median. Pinball loss is strictly consistent for a quantile. In practice, RMSE and MSE reward models that predict the conditional mean, and MAE rewards models that predict the conditional median. When you only need a metric that holds up against outliers, scikit-learn points to median_absolute_error. It takes the median of the absolute residuals instead of their mean.

    Which property of the outcome each loss rewards (from scikit-learn's table of strictly consistent scoring functions)
    Target property of YStrictly consistent lossscikit-learn metric functions
    meansquared errormean_squared_error, root_mean_squared_error
    medianabsolute errormean_absolute_error
    quantilepinball lossmean_pinball_loss

    On Databricks you usually record more than one of these metrics. The MLflow Logged Models example computes RMSE, MAE and R² together from one set of predictions and logs all three against the model. That lets you see whether a model that wins on RMSE also wins on MAE.

    Checkpoint 2 of 5· Check yourself

    A team wants a single error number that is robust to a handful of extreme residuals. Which scikit-learn function is described as robust to outliers?

    Checkpoint 3 of 5· Exam question

    A data scientist is evaluating a regression model that predicts package delivery times. The historical dataset contains a handful of extreme outlier deliveries caused by rare shipping delays, and the team wants an evaluation metric whose value is not dominated by those few extreme residuals. Which metric best fits this requirement?

    Sources43

    3.Relative and worst-case errors: MAPE, MSLE and max_error

    RMSE and MAE are absolute measures: an error of 10 counts the same whether the true value was 20 or 2 million. When targets span several orders of magnitude, that is misleading. scikit-learn's MAPE example uses true values of 1, 10 and 1e6. MAE would have ignored the small values and reflected only the error on the largest one. The mean absolute percentage error (MAPE) divides each error by the true value. This makes it sensitive to relative error and leaves it unchanged by a global scaling of the target.

    Two more metrics answer narrower questions. The mean squared logarithmic error (MSLE) squares the difference of the logs. scikit-learn recommends it for targets with exponential growth, such as population counts. Note that it penalises under-prediction more than over-prediction. Its root version is root_mean_squared_log_error. max_error reports only the single worst residual, the worst-case error between prediction and truth. Neither of these metrics supports every option the core metrics do. For example, max_error and median_absolute_error do not support multioutput.

    Databricks AutoML offers a narrower menu. When you launch a regression run, its primary_metric parameter, which is used to evaluate and rank models, accepts only four values: r2, mae, rmse and mse. To rank on MAPE or MSLE, you compute them yourself.

    Checkpoint 4 of 5· Match them up

    Match each metric to the situation it is designed for

    Tap a term, then the definition that fits it.

    Checkpoint 5 of 5· Check yourself

    Which of these is NOT a supported primary_metric for a Databricks AutoML regression run?

    Sources53

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.MSE and RMSE are in the same units as the target, so either can be read as a typical error size.Why is that wrong?

      MSE is in squared units. Only RMSE, its square root, is back in the target's units.

      Covered in MSE and RMSE: averaging squared residuals

    2. 2.scikit-learn's mean_absolute_percentage_error returns a value from 0 to 100.Why is that wrong?

      It returns a relative value: 200% comes back as 2. Multiply by 100 for the usual percentage.

      Covered in Relative and worst-case errors: MAPE, MSLE and max_error

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 3.
      “RMSE is available through the root_mean_squared_error function.”
      ↩︎ MSE and RMSE: averaging squared residuals
      “mean squared error, a risk metric corresponding to the expected value of the squared (quadratic) error or loss”
      ↩︎ MSE and RMSE: averaging squared residuals
      “The mean_absolute_error function computes mean absolute error”
      ↩︎ MAE, and what each metric is really measuring
      “The median_absolute_error is particularly interesting because it is robust to outliers.”
      ↩︎ MAE, and what each metric is really measuring
      “it would have ignored the small magnitude values and only reflected the error in prediction of highest magnitude value.”
      ↩︎ Relative and worst-case errors: MAPE, MSLE and max_error
      “Note that this metric penalizes an under-predicted estimate greater than an over-predicted estimate.”
      ↩︎ Relative and worst-case errors: MAPE, MSLE and max_error
      “a metric that captures the worst case error between the predicted value and the true value”
      ↩︎ Relative and worst-case errors: MAPE, MSLE and max_error
      “use a strictly consistent scoring function for that (target) functional”
      ↩︎ Key concept
      “another common metric that provides a measure in the same units as the target variable”
      ↩︎ Exam trap 1
      “Thus, an error of 200% corresponds to a relative error of 2.”
      ↩︎ Exam trap 2
      “another common metric that provides a measure in the same units as the target variable”
      ↩︎ Prediction
      “This metric is best to use when targets having exponential growth, such as population counts”
      ↩︎ Checkpoint
    2. 4.
      “(rmse, mae, r2) = compute_metrics(train_y, predictions)”
      ↩︎ MAE, and what each metric is really measuring
    3. 5.
      “Supported metrics for regression: “r2” (default), “mae”, “rmse”, “mse””
      ↩︎ Relative and worst-case errors: MAPE, MSLE and max_error

    Continue to page 2 of 2

    R-squared and choosing a regression metric on Databricks

    Spotted a mistake, or was something unclear? Tell us.