CertSafari
    Snowflake SnowPro Advanced: Data Scientist (DSA-C03)· Lessons

    Domain 1 · Lesson 3/16

    Evaluating, Explaining, Deploying, Monitoring and Versioning ML Models in Snowflake

    Summarize the machine learning lifecycle.

    11 min read
    4.25% of exam
    6 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Calculate and interpret precision, recall and accuracy from a confusion matrix
    • Explain a prediction using Shapley values and background data
    • Promote a model version to production with default versions or aliases in the Snowflake Model Registry
    • Describe how a model version monitor tracks performance and drift over time

    1.Evaluation: accuracy, precision, recall and the confusion matrix

    Evaluating a classifier starts with four counts. True positives (TP) and true negatives (TN) are correct predictions. False positives (FP) are wrong predictions of the positive class, and false negatives (FN) are positive cases the model missed. The main metrics are built from these counts:

    - Accuracy is all correct predictions divided by all predictions, which is (TP + TN) / total. - Precision is TP divided by everything predicted positive, which is TP / (TP + FP). - Recall, also called sensitivity, is TP divided by everything that is actually positive, which is TP / (TP + FN). - F1 is the harmonic mean of precision and recall, and it is useful when the class distribution is uneven.

    The confusion matrix lays all four counts out as a table of actual class against predicted class. The aim is to maximize the counts on the diagonal and minimize the counts off it.

    A binary confusion matrix returned as rows of actual class, predicted class and counttext
    +--------------+--------------+-----------------+-------+------+
    | DATASET_TYPE | ACTUAL_CLASS | PREDICTED_CLASS | COUNT | LOGS |
    |--------------+--------------+-----------------+-------+------|
    | EVAL         | false        | false           |    37 | NULL |
    | EVAL         | false        | true            |     1 | NULL |
    | EVAL         | true         | false           |     0 | NULL |
    | EVAL         | true         | true            |    22 | NULL |
    +--------------+--------------+-----------------+-------+------+

    Read the matrix above: TP = 22, TN = 37, FP = 1 and FN = 0. Recall is 22/22, a perfect 1.0, because the model missed no positives. Precision is 22/23, because one negative was wrongly flagged as positive. Evaluation numbers are only honest when they come from data the model never learned from. Snowflake's classification evaluation holds back a random sample of rows, trains a new model without them, and then scores the held-back rows. If you keep choosing hyperparameters by their score on the test set, that set has become part of training, and the score you report will be inflated.

    Checkpoint 1 of 6· Check yourself

    In the confusion matrix above, which row is the false positive?

    Checkpoint 2 of 6· Exam question

    A team is building a churn model on Snowflake. Which sequence places the lifecycle stages in the correct order, from the first step to the last?

    Sources1

    2.Explainability: why the model predicted what it did

    A good metric tells you that a model works. It does not tell you why. A trained model can be a black box that hides bias, and regulated industries such as finance and healthcare may need evidence that it gets the right answers for the right reasons. The Snowflake Model Registry's explainability function uses Shapley values, which split a prediction into the contribution of each input feature. The contributions are measured against an average prediction, and that average is computed from background data: a representative sample that you supply when you log the model, up to 1,000 rows.

    Shapley contributions for one house price prediction of $250,000 against an average of $100,000
    FeatureValueContribution vs. an average house
    Size2000+$50,000
    LocationBeachside+$75,000
    Bedrooms3+$50,000
    PetsNo-$25,000

    The contributions add up to +$150,000, which is exactly the gap between this prediction and the average house. A contribution can be negative, as the pets feature shows here. A model with explainability enabled has an explain method, and you have to pass it input data, because a Shapley value explains a specific prediction.

    Calling the explain function on the default version of a registered modelpython
    reg = Registry(...)
    mv = reg.get_model("diamond_catboost_explain_enabled").default
    explanations = mv.run(input_data, function_name="explain")

    Checkpoint 3 of 6· Check yourself

    A model version was logged with an old Snowpark ML release that did not support explainability. How do you add explainability to it?

    Sources2

    3.Deployment and versioning with the Model Registry

    Deployment begins with logging the model in the Snowflake Model Registry. The registry is the central record of every model, its metrics and its metadata, whether the model was trained in Snowflake or somewhere else. From the registry you can run inference at scale in a warehouse, or use Model Serving to deploy to Snowpark Container Services. Models are first-class Snowflake objects, so normal role-based access control applies to them.

    Model privileges and what they allow
    PrivilegeAllows
    OWNERSHIPFull control: manage versions, access artifacts, update metadata
    USAGEWarehouse inference and SHOW commands; no access to code, weights or artifacts
    READSPCS inference, model files, metadata and SHOW commands

    Versioning is what makes deployment safe. A model can have unlimited versions, and you log each new one by calling log_model again with the same model_name and a new version_name. A version cannot be changed after it is logged, but you can attach metrics such as test accuracy or a confusion matrix to it so versions can be compared. One version is always the default, and by convention the default is the production version. Promoting a version is a single statement, and production code does not need to change.

    Checkpoint 4 of 6· Fill the gap

    Which keyword promotes new_version to production under the default-version scheme?

    ALTER MODEL my_model SET  ?  = new_version;

    If your organization uses more stages than development and production, such as canary, staging and deprecation, use aliases. An alias points to exactly one version at a time. Every model also has three system aliases: DEFAULT, FIRST (the oldest version) and LAST (the newest version). Tags and separate schemas are two further ways to manage the lifecycle.

    Checkpoint 5 of 6· Match them up

    Match each system alias to the version it refers to

    Tap a term, then the definition that fits it.

    Sources345

    4.Monitoring models in production

    A deployed model gets worse over time without anyone touching it. Inputs drift, the assumptions made at training time go stale, and data pipelines break. ML Observability tracks performance, drift and volume for models deployed through the Model Registry, and it can split those metrics by segment, such as region. It works on stored inference data. Each row in the monitoring log holds an ID, a timestamp, the features, the prediction and a ground truth label, so prediction-based metrics such as accuracy can be compared with what actually happened.

    You create one monitor for each model version with CREATE MODEL MONITOR. A monitor stores data in aggregation windows of at least one day. It can also take an optional baseline table, which drift metrics compare against. Segment columns must be string columns, so numeric values have to be bucketed into categories first. You can also set alerts on performance thresholds. When monitoring shows a model falling below its threshold, the lifecycle loops back: you retrain, log a new version, and promote it.

    Checkpoint 6 of 6· Check yourself

    A team wants one model monitor to watch both v1 and v2 of a churn model during a rollout. What happens?

    Sources6

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.High accuracy proves a classifier is good, even on heavily imbalanced data.Why is that wrong?

      On imbalanced data the majority class drives accuracy. Check precision and recall on the class you care about.

      Covered in Evaluation: accuracy, precision, recall and the confusion matrix

    2. 2.You can register a model and leave it with no production version until it is ready.Why is that wrong?

      A model always has a default version. To block early use, log a placeholder version that throws an error and keep it as the default until a real version is ready.

      Covered in Deployment and versioning with the Model Registry

    3. 3.A single model monitor can cover every version of a model.Why is that wrong?

      Monitors and versions are paired one to one. Every version you want to watch needs its own monitor.

      Covered in Monitoring models in production

    Practise it for real

    Record an evaluation metric on model versions, rank the versions by it, and promote the best one to production

    1. 1.Run: ALTER MODEL my_model MODIFY VERSION v1 SET METADATA = '{"metric": {"accuracy": 0.769}}'; then do the same for each of your other versions with their own scores.

      Why: Metrics stored on each version make the versions comparable in the registry.

      You should see: The statement succeeds, and the version's metadata now includes the accuracy value.

    2. 2.Query my_database.INFORMATION_SCHEMA.MODEL_VERSIONS and select metadata:metric:accuracy AS accuracy, ordering by accuracy DESC.

      Why: The information schema lets you query the registry itself in SQL.

      You should see: One row per model version, with the highest accuracy first.

    3. 3.Run: ALTER MODEL my_model SET DEFAULT_VERSION = <best_version>;

      Why: Production code calls the default version, so this one statement promotes the winning version.

      You should see: SHOW VERSIONS IN MODEL my_model shows the chosen version as the default.

    Stuck? Get a nudge

    You need an existing model with at least two logged versions. If you only have one, log another with log_model, using the same model_name and a new version_name.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Recall (Sensitivity): The ratio of true positives to the total actual positives.”
      ↩︎ Evaluation: accuracy, precision, recall and the confusion matrix
      “The objective is to maximize the number of instances on the diagonal of the matrix while minimizing the number of off-diagonal instances.”
      ↩︎ Evaluation: accuracy, precision, recall and the confusion matrix
      “This metric can be misleading in unbalanced cases.”
      ↩︎ Exam trap 1
      “This metric can be misleading in unbalanced cases.”
      ↩︎ Prediction
      “False Positives (FP): Incorrect predictions of positive instances”
      ↩︎ Checkpoint
    2. 2.
      “Shapley values are a way to attribute the output of a machine learning model to its input features.”
      ↩︎ Explainability: why the model predicted what it did
      “You can provide up to 1,000 rows of background data when logging a model”
      ↩︎ Explainability: why the model predicted what it did
      “Since model versions are immutable, you must create a new model version to add explainability to an existing model.”
      ↩︎ Checkpoint
    3. 3.
      “You can use Model Serving to deploy the models to Snowpark Container Services for inference.”
      ↩︎ Deployment and versioning with the Model Registry
    4. 4.
      “production code only ever calls the default version of the model.”
      ↩︎ Deployment and versioning with the Model Registry
      “A model must always have a default version.”
      ↩︎ Exam trap 2
    5. 5.
      “Metrics are key-value pairs used to track prediction accuracy and other model version characteristics.”
      ↩︎ Deployment and versioning with the Model Registry
      “FIRST refers to the oldest version of the model by creation time.”
      ↩︎ Checkpoint
    6. 6.
      “Model behavior can change over time due to input drift, stale training assumptions, and data pipeline issues”
      ↩︎ Monitoring models in production
      “Segment columns must be string columns in your source data.”
      ↩︎ Monitoring models in production
      “Each model version can have exactly one monitor, and each monitor can monitor exactly one model version; they cannot be shared.”
      ↩︎ Exam trap 3
      “Each model version can have exactly one monitor, and each monitor can monitor exactly one model version; they cannot be shared.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 16 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.