What you will be able to do
- Compute MSE, RMSE and MAE by hand and with scikit-learn, and state the units each one is reported in
- Explain why squared-error metrics target the mean of the response and absolute-error metrics target the median
- Recognise when MAPE, MSLE, median absolute error or max_error answers a question that RMSE or MAE cannot
- Name the regression metrics Databricks AutoML accepts as primary_metric
Key concept
Strictly consistent scoring function — Every regression metric quietly assumes which property of the outcome you are predicting. Squared-error metrics reward predicting the mean, and absolute-error metrics reward predicting the median. So you choose a metric by deciding first what the prediction is supposed to estimate.
1.MSE and RMSE: averaging squared residuals
Every regression metric starts from the residual: the true value y minus the prediction ŷ for each row. The metrics differ in how they turn a column of residuals into one number. The most common choice squares each residual and averages them. That gives the mean squared error (MSE), which scikit-learn describes as the expected value of the squared (quadratic) error. scikit-learn's own small example uses four rows. The residuals are 0.5, −0.5, 0 and −1, so the squared residuals are 0.25, 0.25, 0 and 1, and their average is 0.375.
>>> from sklearn.metrics import mean_squared_error
>>> y_true = [3, -0.5, 2, 7]
>>> y_pred = [2.5, 0.0, 2, 8]
>>> mean_squared_error(y_true, y_pred)
0.375Because MSE is in squared units, an MSE of 0.375 is hard to interpret directly. The root mean squared error (RMSE) is the square root of the MSE: √0.375 ≈ 0.61, in the same units as the target. scikit-learn provides it as root_mean_squared_error. You will also see it computed by hand. A Databricks TabFM tutorial wraps mean_squared_error in np.sqrt, and the SparkR glm tutorial builds RMSE directly from a column of errors. Both are the same formula: square the errors, average them, then take the square root.
# Calculate RMSE
head(select(errors, alias(sqrt(sum(errors$error^2 , na.rm = TRUE) / nrow(errors)), "RMSE")))Checkpoint 1 of 5· Fill the gap
The Databricks TabFM tutorial computes RMSE from scikit-learn's mean_squared_error. Which NumPy function fills the blank?
rmse = np. ? (mean_squared_error(y_test_reg, reg_pred))
r2 = r2_score(y_test_reg, reg_pred)RMSE is the square root of MSE. np.log would give something that is not a standard error metric, and np.abs or np.mean would leave the value in squared units.
Source: docs.databricks.com2.MAE, and what each metric is really measuring
The mean absolute error (MAE) averages the absolute residuals instead of squaring them. On the same four rows, the absolute residuals are 0.5, 0.5, 0 and 1, so the MAE is 0.5. Like RMSE, MAE is in the units of the target. The two numbers differ because squaring weights large residuals more heavily. In the MSE, the single residual of 1 contributes 1 out of a total of 1.5. In the MAE, it contributes 1 out of a total of 2.
scikit-learn's guide on choosing a scoring function answers this. The outcome Y is random for any given set of features, so a point prediction has to estimate some property of its distribution. Squared error is strictly consistent for the mean. Absolute error is strictly consistent for the median. Pinball loss is strictly consistent for a quantile. In practice, RMSE and MSE reward models that predict the conditional mean, and MAE rewards models that predict the conditional median. When you only need a metric that holds up against outliers, scikit-learn points to median_absolute_error. It takes the median of the absolute residuals instead of their mean.
| Target property of Y | Strictly consistent loss | scikit-learn metric functions |
|---|---|---|
| mean | squared error | mean_squared_error, root_mean_squared_error |
| median | absolute error | mean_absolute_error |
| quantile | pinball loss | mean_pinball_loss |
On Databricks you usually record more than one of these metrics. The MLflow Logged Models example computes RMSE, MAE and R² together from one set of predictions and logs all three against the model. That lets you see whether a model that wins on RMSE also wins on MAE.
Checkpoint 2 of 5· Check yourself
A team wants a single error number that is robust to a handful of extreme residuals. Which scikit-learn function is described as robust to outliers?
median_absolute_error takes the median of the absolute residuals, so a few extreme values do not move it much. Squared-error metrics amplify large residuals, and max_error reports only the single worst one.
“The median_absolute_error is particularly interesting because it is robust to outliers.”Source: scikit-learn.org
Checkpoint 3 of 5· Exam question
A data scientist is evaluating a regression model that predicts package delivery times. The historical dataset contains a handful of extreme outlier deliveries caused by rare shipping delays, and the team wants an evaluation metric whose value is not dominated by those few extreme residuals. Which metric best fits this requirement?
Correct answer: A — Mean absolute error, because it averages the absolute value of every residual and weights each error linearly, so a few large residuals do not dominate the score
- A. Mean absolute error sums the absolute residuals and divides by the count, so a single very large error contributes proportionally to its size rather than to its square, making it the metric least distorted by a handful of outliers.
- B. Root mean squared error squares each residual before averaging, so the rare large delivery-time errors contribute disproportionately to the final score, which is the opposite of what the team wants.
- C. R-squared reports explained variance relative to a mean baseline and does not become more robust to outliers as more data is collected; it is still computed from the same squared residuals used in RMSE.
- D. Mean squared error also squares residuals before averaging, so large errors are amplified rather than cancelled, meaning this metric is just as sensitive to the outliers as RMSE.
3.Relative and worst-case errors: MAPE, MSLE and max_error
RMSE and MAE are absolute measures: an error of 10 counts the same whether the true value was 20 or 2 million. When targets span several orders of magnitude, that is misleading. scikit-learn's MAPE example uses true values of 1, 10 and 1e6. MAE would have ignored the small values and reflected only the error on the largest one. The mean absolute percentage error (MAPE) divides each error by the true value. This makes it sensitive to relative error and leaves it unchanged by a global scaling of the target.
Two more metrics answer narrower questions. The mean squared logarithmic error (MSLE) squares the difference of the logs. scikit-learn recommends it for targets with exponential growth, such as population counts. Note that it penalises under-prediction more than over-prediction. Its root version is root_mean_squared_log_error. max_error reports only the single worst residual, the worst-case error between prediction and truth. Neither of these metrics supports every option the core metrics do. For example, max_error and median_absolute_error do not support multioutput.
Databricks AutoML offers a narrower menu. When you launch a regression run, its primary_metric parameter, which is used to evaluate and rank models, accepts only four values: r2, mae, rmse and mse. To rank on MAPE or MSLE, you compute them yourself.
Checkpoint 4 of 5· Match them up
Match each metric to the situation it is designed for
Tap a term, then the definition that fits it.
Each metric changes what counts as a large error: relative to the true value (MAPE), on a log scale (MSLE), only the worst row (max_error), or squared and then returned to the target's units (RMSE).
“This metric is best to use when targets having exponential growth, such as population counts”Source: scikit-learn.org
Checkpoint 5 of 5· Check yourself
Which of these is NOT a supported primary_metric for a Databricks AutoML regression run?
AutoML regression supports r2 (the default), mae, rmse and mse. MAPE exists in scikit-learn, but you cannot use it to rank an AutoML run.
“Supported metrics for regression: “r2” (default), “mae”, “rmse”, “mse””Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.MSE and RMSE are in the same units as the target, so either can be read as a typical error size.Why is that wrong?
MSE is in squared units. Only RMSE, its square root, is back in the target's units.
Covered in MSE and RMSE: averaging squared residuals
2.scikit-learn's mean_absolute_percentage_error returns a value from 0 to 100.Why is that wrong?
It returns a relative value: 200% comes back as 2. Multiply by 100 for the usual percentage.
Covered in Relative and worst-case errors: MAPE, MSLE and max_error
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/ai-runtime/examples/tutorials/sgc-tabfmOfficial docs
“rmse = np.sqrt(mean_squared_error(y_test_reg, reg_pred))”
↩︎ MSE and RMSE: averaging squared residuals - 2.
“# Calculate RMSE”
↩︎ MSE and RMSE: averaging squared residuals - 3.https://scikit-learn.org/stable/modules/model_evaluation.htmlSecondary source
“RMSE is available through the root_mean_squared_error function.”
↩︎ MSE and RMSE: averaging squared residuals“mean squared error, a risk metric corresponding to the expected value of the squared (quadratic) error or loss”
↩︎ MSE and RMSE: averaging squared residuals“The mean_absolute_error function computes mean absolute error”
↩︎ MAE, and what each metric is really measuring“The median_absolute_error is particularly interesting because it is robust to outliers.”
↩︎ MAE, and what each metric is really measuring“it would have ignored the small magnitude values and only reflected the error in prediction of highest magnitude value.”
↩︎ Relative and worst-case errors: MAPE, MSLE and max_error“Note that this metric penalizes an under-predicted estimate greater than an over-predicted estimate.”
↩︎ Relative and worst-case errors: MAPE, MSLE and max_error“a metric that captures the worst case error between the predicted value and the true value”
↩︎ Relative and worst-case errors: MAPE, MSLE and max_error“use a strictly consistent scoring function for that (target) functional”
↩︎ Key concept“another common metric that provides a measure in the same units as the target variable”
↩︎ Exam trap 1“Thus, an error of 200% corresponds to a relative error of 2.”
↩︎ Exam trap 2“another common metric that provides a measure in the same units as the target variable”
↩︎ Prediction“This metric is best to use when targets having exponential growth, such as population counts”
↩︎ Checkpoint - 4.
“(rmse, mae, r2) = compute_metrics(train_y, predictions)”
↩︎ MAE, and what each metric is really measuring - 5.
“Supported metrics for regression: “r2” (default), “mae”, “rmse”, “mse””
↩︎ Relative and worst-case errors: MAPE, MSLE and max_error