CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 39/48

    R-squared and choosing a regression metric on Databricks

    Use common regression metrics: RMSE, MAE, R-squared, etc.

    10 min read
    2.08% of exam
    4 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Interpret an R² value, including 0, 1 and negative scores
    • Explain why R² should not be compared across different datasets and how r2_score handles a constant target
    • Read scikit-learn scorer names such as neg_mean_squared_error correctly
    • Choose a primary regression metric for AutoML and log RMSE, MAE and R² against a model in MLflow

    1.R²: the share of variance a model explains

    RMSE and MAE report error in the target's units, so whether they count as "good" depends on the scale of the problem. The coefficient of determination, R², gives a unitless alternative. It is the proportion of the variance in y that the model explains. scikit-learn computes it with r2_score. Databricks SQL also has a regr_r2 aggregate that returns the coefficient of determination between a dependent and an independent column.

    R² has three reference points. A score of 1.0 means perfect predictions. A score of 0.0 is what a constant model gets if it always predicts the average of y and ignores the features. Scores below 0 mean the model does worse than that constant. On the four-row example (true values 3, −0.5, 2, 7), r2_score returns 0.948. Two caveats matter. First, variance depends on the dataset, so R² may not be meaningfully comparable across different datasets. Second, scikit-learn's r2_score computes unadjusted R².

    There is one edge case to know. If the true target is constant, R² is mathematically NaN for perfect predictions and −Inf for imperfect ones. Either value can break grid-search cross-validation. By default, r2_score replaces them with 1.0 and 0.0. Pass force_finite=False to get the raw definition.

    r2_score on a constant target, with and without force_finitepython
    >>> y_true = [-2, -2, -2]
    >>> y_pred = [-2, -2, -2]
    >>> r2_score(y_true, y_pred)
    1.0
    >>> r2_score(y_true, y_pred, force_finite=False)
    nan

    Checkpoint 1 of 5· Check yourself

    In a validation fold, every true target value is identical and the model's predictions are slightly off. With default settings, what does scikit-learn's r2_score return?

    Sources12

    2.How tools report regression metrics: score() and the neg_ prefix

    You will often meet a regression metric without calling a metric function yourself. scikit-learn evaluates models in three ways. Every estimator has a score method with a default criterion, and for regressors that default is R². Model-selection tools such as GridSearchCV take a scoring parameter. Finally, the sklearn.metrics module provides standalone functions such as mean_squared_error. If you pass scoring=None, the estimator's own score method is used, so a regressor's grid search ranks on R² unless you specify otherwise.

    Every scorer object follows the convention that higher return values are better. Metrics where lower is better are therefore exposed with a neg_ prefix: mean squared error is available as the scorer string 'neg_mean_squared_error', which returns the negated value. A cross-validation result of −0.375 corresponds to an MSE of 0.375. Databricks AutoML handles direction for you. You name the metric used to evaluate and rank models (r2, mae, rmse or mse), and AutoML ranks the trials on it.

    Checkpoint 2 of 5· Check yourself

    You call model.score(X_test, y_test) on a fitted scikit-learn regressor without changing anything. What does it return?

    Checkpoint 3 of 5· Exam question

    After training a regression model to predict house prices, a machine learning engineer evaluates it on a holdout set and finds an R-squared value of -0.15. What does this result indicate about the model?

    Sources32

    3.Choosing the metric, then recording it with the model

    scikit-learn's advice on choosing a metric starts with one rule: if the scoring function is given, for example by a business requirement, use that one. If you are free to choose, decide what the prediction should estimate, and pick a strictly consistent scoring function for it. Use squared error, and therefore RMSE, MSE or R², for the mean. Use absolute error for the median. Ideally, use the same function as the training loss and the evaluation metric. Because R² gives the same ranking as squared error, R², MSE and RMSE always agree on which model is best. MAE can disagree.

    Databricks AutoML regression primary_metric values
    primary_metricMetricUnits / range
    "r2" (default)Coefficient of determinationUnitless; best 1.0, can be negative
    "mae"Mean absolute errorTarget units; lower is better
    "rmse"Root mean squared errorTarget units; lower is better
    "mse"Mean squared errorSquared target units; lower is better
    databricks.automl.regress defaults to ranking runs by R²python
    databricks.automl.regress(
      dataset: Union[pyspark.sql.DataFrame, pandas.DataFrame, pyspark.pandas.DataFrame, str],
      *,
      target_col: str,
      primary_metric: str = "r2",

    Checkpoint 4 of 5· Check yourself

    Model A has lower RMSE than Model B on the same test set. Without computing anything else, which other metric is guaranteed to rank A ahead of B?

    Once you have computed the metrics, record them with the model. In the MLflow 3 Logged Models example, the code predicts, computes RMSE, MAE and R², and passes them to mlflow.log_metrics together with the model_id and the dataset they were measured on. The metrics are then linked to that LoggedModel. You can later filter and order models by metric value on a specific dataset.

    Logging RMSE, R² and MAE against a LoggedModel in MLflowpython
    predictions = lr.predict(train_x)
    (rmse, mae, r2) = compute_metrics(train_y, predictions)
    mlflow.log_metrics(
      metrics={
        "rmse": rmse,
        "r2": r2,
        "mae": mae,
      },
      model_id=logged_model.model_id,
      dataset=train_dataset
    )

    Checkpoint 5 of 5· Exam question

    A team runs a Databricks AutoML regression experiment to forecast equipment repair costs. Large under-predictions are especially costly because they cause the maintenance budget to be badly exceeded, so the team wants the leaderboard's primary metric to penalize large errors more heavily than small ones. Which metric should they configure AutoML to optimize?

    Sources342

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.R² always lies between 0 and 1, so a negative R² must be a bug.Why is that wrong?

      R² has no lower bound. A score below 0 means the model does worse than always predicting the mean of y.

      Covered in R²: the share of variance a model explains

    2. 2.A cross-validation score of −0.375 from 'neg_mean_squared_error' means the model's error is negative.Why is that wrong?

      Scorers negate lower-is-better metrics so that higher is always better. The MSE is 0.375.

      Covered in How tools report regression metrics: score() and the neg_ prefix

    3. 3.A model with R² of 0.9 on one dataset is better than a model with R² of 0.7 on a different dataset.Why is that wrong?

      R² is relative to each dataset's variance, so it may not be meaningfully comparable across datasets.

      Covered in R²: the share of variance a model explains

    Practise it for real

    Reproduce scikit-learn's reference values for MSE, RMSE, MAE and R² on one small dataset, and see the constant-target edge case.

    1. 1.In a notebook, set y_true = [3, -0.5, 2, 7] and y_pred = [2.5, 0.0, 2, 8], then call mean_squared_error(y_true, y_pred).

      Why: This establishes the squared-error baseline that RMSE and R² are built on.

      You should see: 0.375

    2. 2.Call root_mean_squared_error(y_true, y_pred).

      Why: This confirms that RMSE is the square root of the MSE and is in the target's units.

      You should see: About 0.612, the square root of 0.375

    3. 3.Call mean_absolute_error(y_true, y_pred).

      Why: This compares absolute and squared weighting on the same residuals.

      You should see: 0.5

    4. 4.Call r2_score(y_true, y_pred).

      Why: This shows the unitless proportion of explained variance for the same predictions.

      You should see: 0.948

    5. 5.Set y_true = [-2, -2, -2] and y_pred = [-2, -2, -2], then call r2_score twice: once with default settings and once with force_finite=False.

      Why: This shows how scikit-learn keeps a constant target from producing a non-finite score in grid search.

      You should see: 1.0 with defaults, nan with force_finite=False

    Stuck? Get a nudge

    If root_mean_squared_error cannot be imported, wrap mean_squared_error in np.sqrt instead, as the Databricks TabFM tutorial does.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Returns the coefficient of determination from values of a group where xExpr and yExpr are NOT NULL.”
      ↩︎ R²: the share of variance a model explains
    2. 2.
      “It represents the proportion of variance (of y) that has been explained by the independent variables in the model.”
      ↩︎ R²: the share of variance a model explains
      “A constant model that always predicts the expected (average) value of y, disregarding the input features”
      ↩︎ R²: the share of variance a model explains
      “All scorer objects follow the convention that higher return values are better than lower return values.”
      ↩︎ How tools report regression metrics: score() and the neg_ prefix
      “if the scoring function is given, e.g. in a kaggle competition or in a business context, use that one.”
      ↩︎ Choosing the metric, then recording it with the model
      “it is best used for both: as loss function for model training and as metric/score in model evaluation and model comparison”
      ↩︎ Choosing the metric, then recording it with the model
      “Best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse).”
      ↩︎ Exam trap 1
      “are available as ‘neg_mean_squared_error’ which return the negated value of the metric.”
      ↩︎ Exam trap 2
      “may not be meaningfully comparable across different datasets”
      ↩︎ Exam trap 3
      “Best possible score is 1.0 and it can be negative (because the model can be arbitrarily worse).”
      ↩︎ Prediction
      “the default behaviour of r2_score is to replace them with 1.0 (perfect predictions) or 0.0 (imperfect predictions).”
      ↩︎ Checkpoint
      “Most commonly this is accuracy for classifiers and the coefficient of determination”
      ↩︎ Checkpoint
      “R² gives the same ranking as squared error.”
      ↩︎ Checkpoint
    3. 3.
      “Metric used to evaluate and rank model performance.”
      ↩︎ How tools report regression metrics: score() and the neg_ prefix
      “The databricks.automl.regress method configures an AutoML run to train a regression model.”
      ↩︎ Choosing the metric, then recording it with the model

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.