CertSafari
    Snowflake SnowPro Advanced: Data Scientist (DSA-C03)· Lessons

    Domain 3 · Lesson 12/16

    Validating Models in Snowflake: Confusion Matrix, Thresholds, ROC and Regression Metrics

    Validate a data science model.

    16 min read
    6.2% of exam
    8 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Explain how Snowflake's classification function produces evaluation metrics from data it withholds during training
    • Read a SHOW_CONFUSION_MATRIX result and plot it as a heatgrid in Snowsight
    • Use SHOW_THRESHOLD_METRICS counts to build ROC/PR curves, tune a threshold and estimate an expected payout
    • Pick the right regression error metrics (MAE, MSE, RMSE, MAPE) and identify which inputs each one needs
    • Tell apart per-class, global, threshold and confusion-matrix outputs, and know where accuracy misleads

    Key concept

    Holdout evaluation — Snowflake trains a second copy of the model with some rows held back, predicts those held-back rows, and compares the predictions to the true labels. Every validation metric and confusion-matrix cell describes data the model did not see in training.

    1.Where validation numbers come from

    Validation answers one question: how well does the model do on data it has not seen? Snowflake's documentation states the goal directly: metrics measure how accurately a model predicts new data. Before you read any metric, you need to know how Snowflake produced it.

    For a classification model built with CREATE SNOWFLAKE.ML.CLASSIFICATION, Snowflake does not score the model on its own training rows. It takes a random sample from the full dataset and trains a separate model without those rows. It then runs inference on the held-back rows and compares the predicted classes with the actual ones. The model you deploy is still trained on all the data. The evaluation copy exists only to produce honest numbers.

    Evaluation must be turned on. Each evaluation method (SHOW_EVALUATION_METRICS, SHOW_GLOBAL_EVALUATION_METRICS, SHOW_THRESHOLD_METRICS, SHOW_CONFUSION_MATRIX) returns results only for models where evaluation was enabled when the model was created. Each result row is labelled with a dataset_type, which is currently EVAL. That label tells you the numbers come from the held-out sample.

    Forecasting models follow the same idea, adapted for time series. By default the forecasting function evaluates every model it trains with cross-validation. It trains extra models on subsets of the training data and predicts the withheld periods.

    Checkpoint 1 of 9· Put it in order

    Put the steps Snowflake's classification function follows to produce evaluation metrics in order.

    1. 1.Compare the predicted classes with the actual classes to compute metrics
    2. 2.Run inference on the withheld rows
    3. 3.Select a random sample of rows from the entire dataset
    4. 4.Train an additional model without those sampled rows

    Sources123

    2.Reading the confusion matrix

    The confusion matrix is the most basic output of that comparison. SHOW_CONFUSION_MATRIX returns one row for each combination of actual class and predicted class. The count column holds the number of evaluation rows that fall into that combination. The method takes no arguments. Below is the documented output for a binary model:

    SHOW_CONFUSION_MATRIX output for a binary model (evaluation rows only)text
    +--------------+--------------+-----------------+-------+------+
    | DATASET_TYPE | ACTUAL_CLASS | PREDICTED_CLASS | COUNT | LOGS |
    |--------------+--------------+-----------------+-------+------|
    | EVAL         | false        | false           |    37 | NULL |
    | EVAL         | false        | true            |     1 | NULL |
    | EVAL         | true         | false           |     0 | NULL |
    | EVAL         | true         | true            |    22 | NULL |
    +--------------+--------------+-----------------+-------+------+

    Take true as the positive class and read the rows as pairs. Actual true predicted true (22) is a true positive. Actual false predicted false (37) is a true negative. Actual false predicted true (1) is a false positive. Actual true predicted false (0) is a false negative. That makes 60 evaluation rows, 59 of them on the diagonal. The documentation states the goal: maximize the instances on the diagonal and minimize those off it.

    To see this as a chart in Snowsight, run the call, click Chart, and set Chart Type to Heatgrid. Under Data, set Cell values to NONE, Rows to PREDICTED_CLASS and Columns to ACTUAL_CLASS. Note the orientation: predicted classes run down the rows and actual classes across the columns. If you assume the opposite layout, you will swap false positives and false negatives when you read the heatgrid.

    Checkpoint 2 of 9· Check yourself

    In the confusion matrix above, with true as the positive class, how many false positives did the evaluation model produce?

    Sources1

    3.Thresholds, ROC curves and expected payout

    A confusion matrix is a snapshot at one decision threshold. The classifier actually outputs a probability for each class. A row is assigned to a class only if that probability exceeds the threshold. Each threshold gives a different set of true and false positives and negatives. SHOW_THRESHOLD_METRICS exposes all of these. For each class and each threshold, it returns raw counts plus the metrics derived from them. The documentation says this output can be used to plot ROC and PR curves or to tune the threshold. For a multi-class model, each class is treated one-vs-rest: every instance that does not belong to the class is counted as negative.

    Key SHOW_THRESHOLD_METRICS columns and what they mean
    ColumnMeaning
    thresholdThreshold used to generate predictions
    tp / fp / tn / fnRaw counts of true positives, false positives, true negatives and false negatives for the class
    precisionTrue positives divided by all predicted positives
    recall / tprTrue positives divided by all actual positives (also called sensitivity)
    fprShare of actual negatives incorrectly predicted as positive
    accuracyCorrect predictions (positive and negative) divided by all predictions
    supportTrue positives plus false negatives: how often the class actually occurs

    ROC. The tpr and fpr columns at each threshold are the points of the ROC curve. Precision and recall at each threshold give the PR curve. AUC is one of the metrics Snowflake averages across classes. For binary problems, the classification function is trained with an area-under-the-curve loss. Remember that AUC summarises all thresholds at once, while the model is deployed at a single threshold. When positives are rare, a small false-positive *rate* can still mean many false alerts compared with true ones. Check precision at your chosen threshold, not just AUC.

    Expected payout. The Snowflake documentation does not define a payout or profit metric. It provides what you need to compute one yourself: the tp, fp, tn and fn counts at every threshold. The business supplies the value of each outcome. As an illustration with assumed values (not from the documentation): suppose a caught positive earns 100, a false alarm costs 20, and a missed positive costs 100. For the 60-row matrix in the previous section, that gives 22×100 − 1×20 − 0×100 = 2,180, or about 36 per evaluated row. Repeat this for every threshold row and choose the threshold with the best total. That is threshold tuning measured in money instead of accuracy.

    Checkpoint 3 of 9· Check yourself

    You raise a binary model's decision threshold from 0.5 to 0.8. Under Snowflake's definition, when is a row now classified as positive?

    Checkpoint 4 of 9· Exam question

    A bank's fraud model scored 10,000 held-out transactions at a 0.5 threshold, giving TP=80, FP=120, FN=20 and TN=9,780. Each caught fraud nets $450 after recovery, each false alert costs $15 of analyst time, each missed fraud loses $500, and a correct pass is $0. What is the expected net payout on this sample?

    Sources14

    4.Regression problems: error metrics and residuals

    Regression models predict a number, not a class, so there is no confusion matrix. Instead, error metrics summarise the gap between predicted and actual values. Snowflake uses the same family of metrics in several places. A model monitor attached to a regression model accepts RMSE, MAE, MAPE and MSE as performance metrics. Each requires both a prediction_score column and an actual_score column. Without actual values, there is no error to measure. For MSE, RMSE and MAE, the monitor also reports a 95% Wald confidence interval in CI_VALUE.

    A forecasting model's SHOW_EVALUATION_METRICS returns these point metrics from cross-validation: MAE, MAPE, MDA (mean directional accuracy), MSE and SMAPE. It adds interval metrics: COVERAGE_INTERVAL, the share of actual values that fall inside the prediction interval, and the Winkler score. You can also call it with new out-of-sample data to see how the model handles periods it has never seen.

    Cross-validation metrics returned by a forecasting model's SHOW_EVALUATION_METRICStext
    +--------+--------------------------+--------------+--------------------+------+
    | SERIES | ERROR_METRIC             | METRIC_VALUE | STANDARD_DEVIATION | LOGS |
    +--------+--------------------------+--------------+--------------------+------+
    | NULL   | "MAE"                    |         2.49 |                NaN | NULL |
    | NULL   | "MAPE"                   |        0.084 |                NaN | NULL |
    | NULL   | "MDA"                    |         0.99 |                NaN | NULL |
    | NULL   | "MSE"                    |        8.088 |                NaN | NULL |
    | NULL   | "SMAPE"                  |        0.077 |                NaN | NULL |
    | NULL   | "WINKLER_ALPHA=0.05"     |       12.101 |                NaN | NULL |
    | NULL   | "COVERAGE_INTERVAL=0.95" |            1 |                NaN | NULL |
    +--------+--------------------------+--------------+--------------------+------+

    The STANDARD_DEVIATION column is empty here for a documented reason. With the default n_splits of 1, only one validation set exists, so there is no spread to report. Increase n_splits if you want to see how stable each metric is across folds.

    With a Snowpark ML regressor, you validate the same way you would in scikit-learn. Split the data, fit on the training portion, predict the held-out portion, and score it. The Snowflake developer guide splits with random_split([0.85, 0.15], seed=42), trains an XGBRegressor, and then computes test MSE and R² on the predictions:

    Scoring a Snowpark ML XGBRegressor on a held-out test set with MSE and R²python
    from sklearn.metrics import mean_squared_error, r2_score
    
    predictions = regressor.predict(test_df)
    predictions_pd = predictions.to_pandas()
    
    mse = mean_squared_error(predictions_pd["ETA_MINUTES"], predictions_pd["predicted_eta"])
    r2 = r2_score(predictions_pd["ETA_MINUTES"], predictions_pd["predicted_eta"])
    
    print(f"Test MSE: {mse:.2f}")
    print(f"Test R²: {r2:.2f}")

    Checkpoint 5 of 9· Check yourself

    You want to get RMSE for a regression model from its model monitor. Which columns must the monitor have?

    Checkpoint 6 of 9· Exam question

    A churn team plots the ROC curve for a gradient-boosted model. A missed churner costs the business about ten times more than an unnecessary retention offer, and about 8% of customers churn. How should the operating threshold be chosen from the ROC curve?

    Sources562

    5.Model metrics: which output answers which question

    The classification evaluation APIs split the work across four methods. Choosing the right one is often what an exam question is really testing. Per-class precision, recall and F1 come from SHOW_EVALUATION_METRICS. One overall number per metric comes from SHOW_GLOBAL_EVALUATION_METRICS, which averages the per-class values. Its average_type column is currently MACRO, so every class counts equally regardless of its size. Logistic loss is calculated once for the whole model. F1 is the harmonic mean of precision and recall. The documentation recommends it as the balancing metric when class distributions are uneven.

    Classification evaluation methods compared
    MethodGranularityUse it to
    SHOW_EVALUATION_METRICSOne row per class per error_metricCompare precision, recall and F1 across classes
    SHOW_GLOBAL_EVALUATION_METRICSWhole model; average_type currently MACROReport one headline number per metric
    SHOW_THRESHOLD_METRICSPer class, per threshold, with tp/fp/tn/fnPlot ROC and PR curves; tune the threshold
    SHOW_CONFUSION_MATRIXPer actual_class × predicted_class countPlot a heatgrid; see which classes are confused

    After deployment, the same metric families continue in model monitoring. For a binary classification model, MODEL_MONITOR_PERFORMANCE_METRIC accepts ROC_AUC, CLASSIFICATION_ACCURACY, PRECISION, RECALL and F1_SCORE. For a multi-class model, it accepts accuracy plus macro- and micro-averaged precision and recall. ROC_AUC needs prediction_score and actual_class, because AUC is built from scores. The class-based metrics need prediction_class and actual_class. Do not count support as a quality score. It describes how often a class appears in the dataset, not how the model performs.

    Checkpoint 7 of 9· Match them up

    Match each threshold-metrics column to its meaning.

    Tap a term, then the definition that fits it.

    Checkpoint 8 of 9· Fill the gap

    You need a single macro-averaged precision, recall and F1 for the whole model. Which method completes the second call?

    CALL <model_name>!SHOW_EVALUATION_METRICS();
    CALL <model_name>! ? ();
    CALL <model_name>!SHOW_THRESHOLD_METRICS();
    CALL <model_name>!SHOW_CONFUSION_MATRIX();

    Checkpoint 9 of 9· Exam question

    A hospital builds a screening model for a serious disease that treatment can cure if caught early. A follow-up test for a false alarm is cheap, but a missed case is catastrophic. Which metric should the data scientist prioritize when validating candidate models?

    Sources715

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.You can call SHOW_CONFUSION_MATRIX or SHOW_THRESHOLD_METRICS on any classification model after training.Why is that wrong?

      These methods return results only for models where evaluation was enabled when the model was created. The metrics come from a separately trained evaluation model scored on withheld rows.

      Covered in Where validation numbers come from

    2. 2.High accuracy, or a high ROC-AUC, on a rare-positive problem means the model's alerts are mostly correct.Why is that wrong?

      Accuracy is dominated by the majority class and can mislead on unbalanced data. Check precision and the false-positive count at the deployed threshold, which SHOW_THRESHOLD_METRICS reports per threshold.

      Covered in Thresholds, ROC curves and expected payout

    3. 3.A class with high support is a class the model predicts well.Why is that wrong?

      Support only counts how often the class actually occurs (true positives plus false negatives). It says nothing about prediction quality.

      Covered in Model metrics: which output answers which question

    4. 4.A model monitor can report RMSE or MAE from predictions alone.Why is that wrong?

      Regression error metrics compare predictions with actual values. The monitor needs both prediction_score and actual_score, and actual columns are required for accuracy metrics.

      Covered in Regression problems: error metrics and residuals

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Metrics measure how accurately a model predicts new data.”
      ↩︎ Where validation numbers come from
      “The Snowflake classification currently evaluates models by selecting a random sample from the entire dataset.”
      ↩︎ Where validation numbers come from
      “The objective is to maximize the number of instances on the diagonal of the matrix while minimizing the number of off-diagonal instances.”
      ↩︎ Reading the confusion matrix
      “for Rows select PREDICTED_CLASS, and for Columns select ACTUAL_CLASS”
      ↩︎ Reading the confusion matrix
      “This can be used to plot ROC and PR curves or do threshold tuning if desired.”
      ↩︎ Thresholds, ROC curves and expected payout
      “For binary classification, the model is trained using an area-under-the-curve loss function.”
      ↩︎ Thresholds, ROC curves and expected payout
      “The harmonic mean of precision and recall. It provides a balance between precision and recall, especially when there is an uneven class distribution.”
      ↩︎ Model metrics: which output answers which question
      “an additional model is trained on the original data but with some data points withheld”
      ↩︎ Key concept
      “This metric can be misleading in unbalanced cases.”
      ↩︎ Exam trap 2
      “Support is not itself a metric of the model but a characteristic of the dataset.”
      ↩︎ Exam trap 3
      “The withheld data points are then used for inference, and the predicted classes are compared to the actual classes.”
      ↩︎ Checkpoint
      “This metric can be misleading in unbalanced cases.”
      ↩︎ Prediction
      “The sample is classified as belonging to a class if the predicted probability of being in that class exceeds the specified threshold.”
      ↩︎ Checkpoint
      “Support is not itself a metric of the model but a characteristic of the dataset.”
      ↩︎ Checkpoint
    2. 2.
      “By default, the forecasting function evaluates all models it trains using a method called cross-validation.”
      ↩︎ Where validation numbers come from
      “When n_splits is 1 (the default), the standard deviation for evaluation metric values is NULL, as only a validation dataset is used.”
      ↩︎ Regression problems: error metrics and residuals
    3. 3.
      “The name of the dataset used for metrics calculation, currently EVAL.”
      ↩︎ Where validation numbers come from
      “in models where evaluation was enabled at instantiation”
      ↩︎ Exam trap 1
      “The number of instances of the given combination of actual and predicted class.”
      ↩︎ Checkpoint
    4. 4.
      “Precision for the given class. The ratio of true positives to the total predicted positives.”
      ↩︎ Thresholds, ROC curves and expected payout
    5. 5.
      “Valid values if the model monitor is attached to a regression model: 'RMSE' 'MAE' 'MAPE' 'MSE'”
      ↩︎ Regression problems: error metrics and residuals
      “Wald CI for MSE, RMSE, and MAE.”
      ↩︎ Regression problems: error metrics and residuals
      “ROC_AUC: Requires the prediction_score and actual_class columns”
      ↩︎ Model metrics: which output answers which question
      “RMSE: Requires the prediction_score and actual_score columns”
      ↩︎ Checkpoint
    6. 6.
      “MAE: Mean Absolute Error. MAPE: Mean Absolute Percentage Error. MDA: Mean Directional Accuracy. MSE: Mean Squared Error.”
      ↩︎ Regression problems: error metrics and residuals
    7. 7.
      “The method of aggregation used to calculate overall metrics from the individual class metrics, currently MACRO.”
      ↩︎ Model metrics: which output answers which question

    Also cited

    Ready to test yourself?

    Practise the 22 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.