What you will be able to do
- Interpret an R² value, including 0, 1 and negative scores
- Explain why R² should not be compared across different datasets and how r2_score handles a constant target
- Read scikit-learn scorer names such as neg_mean_squared_error correctly
- Choose a primary regression metric for AutoML and log RMSE, MAE and R² against a model in MLflow
1.R²: the share of variance a model explains
RMSE and MAE report error in the target's units, so whether they count as "good" depends on the scale of the problem. The coefficient of determination, R², gives a unitless alternative. It is the proportion of the variance in y that the model explains. scikit-learn computes it with r2_score. Databricks SQL also has a regr_r2 aggregate that returns the coefficient of determination between a dependent and an independent column.
R² has three reference points. A score of 1.0 means perfect predictions. A score of 0.0 is what a constant model gets if it always predicts the average of y and ignores the features. Scores below 0 mean the model does worse than that constant. On the four-row example (true values 3, −0.5, 2, 7), r2_score returns 0.948. Two caveats matter. First, variance depends on the dataset, so R² may not be meaningfully comparable across different datasets. Second, scikit-learn's r2_score computes unadjusted R².
There is one edge case to know. If the true target is constant, R² is mathematically NaN for perfect predictions and −Inf for imperfect ones. Either value can break grid-search cross-validation. By default, r2_score replaces them with 1.0 and 0.0. Pass force_finite=False to get the raw definition.
>>> y_true = [-2, -2, -2]
>>> y_pred = [-2, -2, -2]
>>> r2_score(y_true, y_pred)
1.0
>>> r2_score(y_true, y_pred, force_finite=False)
nanCheckpoint 1 of 5· Check yourself
In a validation fold, every true target value is identical and the model's predictions are slightly off. With default settings, what does scikit-learn's r2_score return?
By default (force_finite=True), the non-finite −Inf for imperfect predictions is replaced with 0.0, so grid search can continue. With force_finite=False, you would get -inf.
“the default behaviour of r2_score is to replace them with 1.0 (perfect predictions) or 0.0 (imperfect predictions).”Source: scikit-learn.org
2.How tools report regression metrics: score() and the neg_ prefix
You will often meet a regression metric without calling a metric function yourself. scikit-learn evaluates models in three ways. Every estimator has a score method with a default criterion, and for regressors that default is R². Model-selection tools such as GridSearchCV take a scoring parameter. Finally, the sklearn.metrics module provides standalone functions such as mean_squared_error. If you pass scoring=None, the estimator's own score method is used, so a regressor's grid search ranks on R² unless you specify otherwise.
Every scorer object follows the convention that higher return values are better. Metrics where lower is better are therefore exposed with a neg_ prefix: mean squared error is available as the scorer string 'neg_mean_squared_error', which returns the negated value. A cross-validation result of −0.375 corresponds to an MSE of 0.375. Databricks AutoML handles direction for you. You name the metric used to evaluate and rank models (r2, mae, rmse or mse), and AutoML ranks the trials on it.
First negate the score to get the MSE, which is 4.0. Then take the square root: the RMSE is 2.0, in the target's units.
Checkpoint 2 of 5· Check yourself
You call model.score(X_test, y_test) on a fitted scikit-learn regressor without changing anything. What does it return?
The default score method returns accuracy for classifiers and R² for regressors. Negated error metrics appear only when you request them by scorer name.
“Most commonly this is accuracy for classifiers and the coefficient of determination”Source: scikit-learn.org
Checkpoint 3 of 5· Exam question
After training a regression model to predict house prices, a machine learning engineer evaluates it on a holdout set and finds an R-squared value of -0.15. What does this result indicate about the model?
Correct answer: A — The model fits the holdout data worse than a naive baseline that always predicts the average house price, so the learned relationship is actively misleading rather than simply weak
- A. R-squared compares the model's squared error against the squared error of always predicting the mean; a negative value means the fitted model performs worse than that trivial baseline, which is a genuine warning sign rather than a rounding artifact.
- B. R-squared is not a percentage scaled from zero to one hundred in this way, and a negative value does not mean the model is merely fifteen percent worse than perfect; it means the model underperforms the mean-prediction baseline.
- C. The coefficient of determination is only bounded above by one; it is unbounded below and can legitimately go negative whenever a model fits worse than the mean, so a negative value alone does not indicate a pipeline bug.
- D. Root mean squared error and mean absolute error are unrelated in magnitude to whether R-squared is negative, and neither metric is forced to zero by a negative coefficient of determination.
3.Choosing the metric, then recording it with the model
scikit-learn's advice on choosing a metric starts with one rule: if the scoring function is given, for example by a business requirement, use that one. If you are free to choose, decide what the prediction should estimate, and pick a strictly consistent scoring function for it. Use squared error, and therefore RMSE, MSE or R², for the mean. Use absolute error for the median. Ideally, use the same function as the training loss and the evaluation metric. Because R² gives the same ranking as squared error, R², MSE and RMSE always agree on which model is best. MAE can disagree.
| primary_metric | Metric | Units / range |
|---|---|---|
| "r2" (default) | Coefficient of determination | Unitless; best 1.0, can be negative |
| "mae" | Mean absolute error | Target units; lower is better |
| "rmse" | Root mean squared error | Target units; lower is better |
| "mse" | Mean squared error | Squared target units; lower is better |
databricks.automl.regress(
dataset: Union[pyspark.sql.DataFrame, pandas.DataFrame, pyspark.pandas.DataFrame, str],
*,
target_col: str,
primary_metric: str = "r2",Checkpoint 4 of 5· Check yourself
Model A has lower RMSE than Model B on the same test set. Without computing anything else, which other metric is guaranteed to rank A ahead of B?
R² is a rescaling of squared error on a fixed dataset, so it ranks models exactly as MSE and RMSE do. MAE, MAPE and median absolute error weight residuals differently and can disagree.
“R² gives the same ranking as squared error.”Source: scikit-learn.org
Once you have computed the metrics, record them with the model. In the MLflow 3 Logged Models example, the code predicts, computes RMSE, MAE and R², and passes them to mlflow.log_metrics together with the model_id and the dataset they were measured on. The metrics are then linked to that LoggedModel. You can later filter and order models by metric value on a specific dataset.
predictions = lr.predict(train_x)
(rmse, mae, r2) = compute_metrics(train_y, predictions)
mlflow.log_metrics(
metrics={
"rmse": rmse,
"r2": r2,
"mae": mae,
},
model_id=logged_model.model_id,
dataset=train_dataset
)Checkpoint 5 of 5· Exam question
A team runs a Databricks AutoML regression experiment to forecast equipment repair costs. Large under-predictions are especially costly because they cause the maintenance budget to be badly exceeded, so the team wants the leaderboard's primary metric to penalize large errors more heavily than small ones. Which metric should they configure AutoML to optimize?
Correct answer: A — Root mean squared error, because squaring each residual before averaging gives disproportionately more weight to large prediction misses than to small ones
- A. Because residuals are squared before being averaged, a repair-cost prediction that misses by a large margin contributes far more to root mean squared error than several small misses, matching the team's goal of penalizing big errors more heavily.
- B. Mean absolute error weights every residual linearly by its magnitude, so a large miss and several small misses of the same total size contribute equally, which does not emphasize large errors the way the team wants.
- C. A higher R-squared reflects more explained variance overall but does not specifically penalize large individual residuals more than small ones, so it is not the right lever for this requirement.
- D. Explained variance summarizes how much of the target's spread the model accounts for but, like R-squared, does not weight large individual residuals more heavily than small ones.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.R² always lies between 0 and 1, so a negative R² must be a bug.Why is that wrong?
R² has no lower bound. A score below 0 means the model does worse than always predicting the mean of y.
Covered in R²: the share of variance a model explains
2.A cross-validation score of −0.375 from 'neg_mean_squared_error' means the model's error is negative.Why is that wrong?
Scorers negate lower-is-better metrics so that higher is always better. The MSE is 0.375.
Covered in How tools report regression metrics: score() and the neg_ prefix
3.A model with R² of 0.9 on one dataset is better than a model with R² of 0.7 on a different dataset.Why is that wrong?
R² is relative to each dataset's variance, so it may not be meaningfully comparable across datasets.
Covered in R²: the share of variance a model explains
Practise it for real
Reproduce scikit-learn's reference values for MSE, RMSE, MAE and R² on one small dataset, and see the constant-target edge case.
1.In a notebook, set y_true = [3, -0.5, 2, 7] and y_pred = [2.5, 0.0, 2, 8], then call mean_squared_error(y_true, y_pred).
Why: This establishes the squared-error baseline that RMSE and R² are built on.
You should see: 0.375
2.Call root_mean_squared_error(y_true, y_pred).
Why: This confirms that RMSE is the square root of the MSE and is in the target's units.
You should see: About 0.612, the square root of 0.375
3.Call mean_absolute_error(y_true, y_pred).
Why: This compares absolute and squared weighting on the same residuals.
You should see: 0.5
4.Call r2_score(y_true, y_pred).
Why: This shows the unitless proportion of explained variance for the same predictions.
You should see: 0.948
5.Set y_true = [-2, -2, -2] and y_pred = [-2, -2, -2], then call r2_score twice: once with default settings and once with force_finite=False.
Why: This shows how scikit-learn keeps a constant target from producing a non-finite score in grid search.
You should see: 1.0 with defaults, nan with force_finite=False
Stuck? Get a nudge
If root_mean_squared_error cannot be imported, wrap mean_squared_error in np.sqrt instead, as the Databricks TabFM tutorial does.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Returns the coefficient of determination from values of a group where xExpr and yExpr are NOT NULL.”
↩︎ R²: the share of variance a model explains - 2.https://scikit-learn.org/stable/modules/model_evaluation.htmlSecondary source
“It represents the proportion of variance (of y) that has been explained by the independent variables in the model.”
↩︎ R²: the share of variance a model explains“A constant model that always predicts the expected (average) value of y, disregarding the input features”
↩︎ R²: the share of variance a model explains“All scorer objects follow the convention that higher return values are better than lower return values.”
↩︎ How tools report regression metrics: score() and the neg_ prefix“if the scoring function is given, e.g. in a kaggle competition or in a business context, use that one.”
↩︎ Choosing the metric, then recording it with the model“it is best used for both: as loss function for model training and as metric/score in model evaluation and model comparison”
↩︎ Choosing the metric, then recording it with the model“Best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse).”
↩︎ Exam trap 1“are available as ‘neg_mean_squared_error’ which return the negated value of the metric.”
↩︎ Exam trap 2“may not be meaningfully comparable across different datasets”
↩︎ Exam trap 3“Best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse).”
↩︎ Prediction“the default behaviour of r2_score is to replace them with 1.0 (perfect predictions) or 0.0 (imperfect predictions).”
↩︎ Checkpoint“Most commonly this is accuracy for classifiers and the coefficient of determination”
↩︎ Checkpoint“R² gives the same ranking as squared error.”
↩︎ Checkpoint - 3.
“Metric used to evaluate and rank model performance.”
↩︎ How tools report regression metrics: score() and the neg_ prefix“The databricks.automl.regress method configures an AutoML run to train a regression model.”
↩︎ Choosing the metric, then recording it with the model - 4.
“These metrics are now linked to the LoggedModel entity”
↩︎ Choosing the metric, then recording it with the model