What you will be able to do
- Create a model version monitor and explain what each required and optional parameter makes possible
- Explain why a monitor with no baseline cannot report drift, and why one with no actuals cannot report accuracy
- Get drift, statistical and performance metrics from a monitor in SQL so you can compare them against thresholds
- Recover a monitor that has suspended itself after refresh failures
- Choose between monitor drift and statistical metrics and Snowflake anomaly detection, and run either on a schedule with a Snowflake alert
Key concept
Model version monitor — A schema-level object that belongs to exactly one model version. It reads that version's stored inference data from a source table or view, groups it into time windows, and computes drift, performance and volume metrics. You can then view those metrics in Snowsight or query them in SQL.
1.What a model monitor watches, and how you create one
A production model can get worse even when no one changes its code. Input data drifts away from what the model was trained on, assumptions made at training time go stale, and upstream pipelines break. ML Observability tracks models deployed through the Snowflake Model Registry across performance, drift and volume. It does this from stored inference data: rows that hold the features, the prediction, a timestamp and, when labels become available, the ground truth. Because it reads stored data, it works the same way whether inference ran inside Snowflake or outside Snowflake with the results written back.
This page covers model version monitors. Snowflake also has gateway model monitors, which watch real-time inference behind a Snowflake Gateway, including A/B tests. You choose the type in the statement: VERSION creates a version monitor and GATEWAY creates a gateway monitor, and one statement cannot specify both. Every model version you want to watch needs its own monitor, created explicitly. A monitor is never shared between versions. Each monitor records the model version it watches, the table that holds its logs, the aggregation window (at least 1 day for version monitors) and an optional baseline.
CREATE [ OR REPLACE ] MODEL MONITOR [ IF NOT EXISTS ] <monitor_name> WITH
MODEL = <model_name>
VERSION = '<version_name>'
FUNCTION = '<function_name>'
SOURCE = <source_name>
WAREHOUSE = <warehouse_name>
REFRESH_INTERVAL = '<num> { seconds | minutes | hours | days }'
AGGREGATION_WINDOW = '<num> days'
TIMESTAMP_COLUMN = <timestamp_name>
[ BASELINE = <baseline_name> ]
[ ID_COLUMNS = <id_column_name_array> ]
[ PREDICTION_CLASS_COLUMNS = <prediction_class_column_name_array> ]
[ PREDICTION_SCORE_COLUMNS = <prediction_column-name_array> ]
[ ACTUAL_CLASS_COLUMNS = <actual_class_column_name_array> ]
[ ACTUAL_SCORE_COLUMNS = <actual_column_name_array> ]
[ SEGMENT_COLUMNS = <segment_column_name_array> ]
[ CUSTOM_METRIC_COLUMNS = <custom_metric_column_name_array> ]
[ COMMENT = '<string_literal>' ]The monitor uses WAREHOUSE for its internal compute. REFRESH_INTERVAL sets how often it refreshes, with a minimum of 60 seconds. TIMESTAMP_COLUMN must be of type TIMESTAMP_NTZ. To create a monitor you need the CREATE MODEL MONITOR privilege on the schema, SELECT on the source, and USAGE on the warehouse and the model. An account can have at most 250 model monitors.
Plan the configuration before you create the monitor. The model and source table you pick cannot be changed afterwards. ALTER MODEL MONITOR changes only a few options, such as segments and suspend/resume. To change anything else, you drop the monitor and create a new one.
Checkpoint 1 of 6· Check yourself
A team wants to point an existing model version monitor at a new source view that adds two feature columns. What do they have to do?
The model and source table are fixed when the monitor is created, so a new source means dropping the monitor and creating a new one. You also cannot add a second monitor, because a version can have only one.
“cannot be changed after the monitor is created. You can modify only a few options using ALTER MODEL MONITOR.”Source: docs.snowflake.com
Tracking data drift and statistical anomalies with Snowflake's integrated monitoring and alerting. Two built-in routes cover this, and they answer different questions.
- Model monitor metrics. Every model version monitor computes drift metrics (distribution changes or data shifts), performance metrics, and statistical metrics (counts or null values). Drift and statistical metrics together show whether the inputs and predictions of a deployed model are moving, or whether the data pipeline is producing unusual counts or nulls. Drift needs a baseline. You read the metrics with the monitor metric functions in SQL.
- Anomaly Detection ML functions. To flag unusual values in a data stream, such as sales falling outside an expected range, you can train a SNOWFLAKE.ML.ANOMALY_DETECTION model and call its DETECT_ANOMALIES method. The rows it returns carry an is_anomaly flag, so you can filter on is_anomaly = TRUE. You can automate this with Snowflake Tasks or Alerts.
In both routes, Snowflake alerts supply the alerting. An alert is a schema-level object with a condition, an action, and a schedule. The action can send an email notification. An alert on a schedule evaluates the condition against the existing data at each run, so it suits a query over monitor metrics or a stored procedure that returns anomalous rows. If the alert condition returns one or more rows, the action runs. The exam point: monitor metrics tell you the model's data has shifted, anomaly detection tells you individual values fall outside the expected range, and alerts are what turn either result into a notification.
Checkpoint 2 of 6· Check yourself
A team wants an email whenever newly scored sales values fall outside the range a trained Snowflake anomaly detection model expects. Which design uses Snowflake's integrated tooling?
Anomaly detection results can be monitored with a Snowflake alert whose condition checks for rows from the detection call and whose action sends the notification. A monitor without a baseline cannot report drift at all.
“monitoring your data for anomalies, by using Anomaly Detection functions within Snowflake Tasks or Alerts”Source: docs.snowflake.com
2.Baselines, actuals and the columns that enable each metric
Most of the optional parameters each switch on one kind of metric, and leaving one out is a common reason a metric is missing.
- Drift compares current data against a reference, so it needs BASELINE: a table holding a snapshot of data shaped like the source. A snapshot of it is embedded in the monitor.
- Accuracy compares predictions with real outcomes. A version monitor needs at least one prediction column. Actual columns are optional, but without them no accuracy metrics are computed. A gateway monitor gets its labels from a GROUND_TRUTH table and computes performance metrics only when ID_COLUMNS and GROUND_TRUTH are both specified.
- Custom metrics: CUSTOM_METRIC_COLUMNS names numeric columns to track that are not treated as model features.
- Segments: SEGMENT_COLUMNS tracks quality per subset, such as region. Segment columns must be STRING, with at most 5 per monitor. If you want to segment by a numeric value, bucket it into categories first.
Column types depend on the model task. For binary classification, predictions can be scores or classes but actuals must be classes. For multi-class classification, both must be classes. For regression, both must be numbers. Each prediction and actual array can hold at most one element, and a column can appear in only one parameter.
| Parameter | What it enables | What happens without it |
|---|---|---|
| BASELINE | Drift metrics against a reference snapshot | The monitor cannot detect drift |
| ACTUAL_CLASS_COLUMNS / ACTUAL_SCORE_COLUMNS | Accuracy metrics for a model version monitor | Accuracy metrics are not computed |
| ID_COLUMNS + GROUND_TRUTH | Performance metrics for a gateway model monitor | Specify both together or omit both |
| CUSTOM_METRIC_COLUMNS | Tracking numeric columns that are not treated as features | Those columns are not tracked as custom metrics |
| SEGMENT_COLUMNS | Metrics for each STRING segment, up to 5 columns | Metrics cover only the complete dataset |
Checkpoint 3 of 6· Match them up
Match each requirement to the parameter that meets it
Tap a term, then the definition that fits it.
Each optional parameter enables one kind of metric. Custom metric columns are tracked without being treated as features, while segments split metrics by STRING columns.
“These columns are not treated as features.”Source: docs.snowflake.com
Checkpoint 4 of 6· Exam question
A team created a model version monitor without a `BASELINE` table. The dashboard shows volume and performance metrics but no drift. They now want drift metrics on this same monitor. What must they do?
Correct answer: A — Drop the monitor and recreate it with `BASELINE` set to a snapshot table of training features and predictions.
- A. Correct: drift needs baseline data, and the baseline cannot be added to an existing monitor, so the monitor must be dropped and recreated with a baseline table.
- B. Incorrect: the basic configuration of a model monitor, including its baseline, cannot be changed after creation, so there is no property to set with ALTER.
- C. Incorrect: the monitor does not infer a baseline from a flag column or from the earliest window; the baseline is a separate table passed at creation.
- D. Incorrect: privileges on a training table do not cause automatic baseline discovery; the baseline must be named explicitly when the monitor is created.
Sources2
3.Querying drift, stat and performance metrics for thresholds
To review metrics by hand, open AI & ML » Models in Snowsight, choose a model and open one of its monitors. The dashboard shows graphs that depend on the model type. To check metrics against thresholds automatically, you need them in SQL. Three table functions return them: MODEL_MONITOR_DRIFT_METRIC, MODEL_MONITOR_PERFORMANCE_METRIC and MODEL_MONITOR_STAT_METRIC. Each one takes the monitor's name and the name of the metric you want.
MODEL_MONITOR_DRIFT_METRIC(
<model_monitor_name>, <drift_metric_name>, <column_name>
[, <granularity> [, <start_time> [, <end_time> [, <extra_args> ] ] ] ]
)The drift metric names are JENSEN_SHANNON, DIFFERENCE_OF_MEANS, WASSERSTEIN and POPULATION_STABILITY_INDEX (PSI). On a version monitor you can compute drift for any feature, prediction or actual column. The stat function finds problems that accuracy can miss. COUNT_NULL shows a sudden rise in nulls, and COUNT shows a drop in daily volume. Both work on every column type. MIN, MAX, AVG and SUM work on numeric columns only, and the default granularity for version monitors is 1 DAY.
The sources do not list the metric names that MODEL_MONITOR_PERFORMANCE_METRIC accepts. What they do establish is that a monitor without actuals has no performance metrics to return. In practice, a threshold is a comparison you write against these function outputs, for example PSI above a limit you choose. The next lesson on alerts shows how to run that comparison on a schedule.
A monitor that cannot refresh also leaves gaps in this data. After five consecutive refresh failures related to the source tables, a monitor suspends itself. When that happens, DESCRIBE MODEL MONITOR shows SUSPENDED in aggregation_status and the SQL error in aggregation_last_error.
Checkpoint 5 of 6· Check yourself
Accuracy looks stable, but data engineers want early warning when null counts spike or daily row volume halves. Which function gives them that?
Stat metrics count rows and nulls for each column in each time window, so they show pipeline breakage directly. Accuracy can stay flat while the data underneath it degrades.
“'COUNT' and 'COUNT_NULL' are supported for all column types.”Source: docs.snowflake.com
Checkpoint 6 of 6· Put it in order
A monitor's source table was renamed and the dashboard has stopped updating. Put the recovery steps in order.
- 1.Run DESCRIBE MODEL MONITOR and read aggregation_status and aggregation_last_error
- 2.The monitor suspends refreshes after five consecutive source-related failures
- 3.Run ALTER MODEL MONITOR … RESUME
- 4.Resolve the root cause of the refresh failure
Suspension happens automatically. DESCRIBE shows the cause, and you resume the monitor only after fixing that cause.
“After resolving the root cause of the refresh failure, resume the monitor by issuing ALTER MODEL MONITOR … RESUME.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A monitor with prediction and actual columns will report drift automatically.Why is that wrong?
Drift needs a BASELINE table. Without one, the monitor computes accuracy from the actuals but cannot detect drift.
Covered in Baselines, actuals and the columns that enable each metric
2.A numeric column such as temperature can be used directly as a segment column.Why is that wrong?
Segment columns must be STRING. Bucket a numeric value into categories, such as COLD, MODERATE and HOT, before you create the monitor.
Covered in Baselines, actuals and the columns that enable each metric
3.You can point an existing monitor at a different source table with ALTER MODEL MONITOR.Why is that wrong?
The model and source cannot be changed after creation. ALTER changes only a few options, so you drop the monitor and create a new one.
Covered in What a model monitor watches, and how you create one
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/model-observabilityOfficial docs
“The minimum time granularity at which data is stored (aggregation window), currently 1 day minimum.”
↩︎ What a model monitor watches, and how you create one“You can create a maximum of 250 model monitors per account.”
↩︎ What a model monitor watches, and how you create one“Drift metrics: Distribution changes or data shifts”
↩︎ What a model monitor watches, and how you create one“You can set up alerts and notifications for your monitoring metrics.”
↩︎ What a model monitor watches, and how you create one“Model monitors automatically suspend refreshes when they encounter five consecutive refresh failures related to the source tables.”
↩︎ Querying drift, stat and performance metrics for thresholds“Each model version can have exactly one monitor, and each monitor can monitor exactly one model version; they cannot be shared.”
↩︎ Key concept“Segment columns must be string columns in your source data.”
↩︎ Exam trap 2“After resolving the root cause of the refresh failure, resume the monitor by issuing ALTER MODEL MONITOR … RESUME.”
↩︎ Checkpoint - 2.
“You cannot specify both VERSION and GATEWAY in the same statement; specify one to indicate the monitor type.”
↩︎ What a model monitor watches, and how you create one“Actual columns are optional, but accuracy metrics are not computed if they are not specified.”
↩︎ Baselines, actuals and the columns that enable each metric“specify ID_COLUMNS and GROUND_TRUTH together to enable performance metrics monitoring, or omit both.”
↩︎ Baselines, actuals and the columns that enable each metric“Although this parameter is optional, if it is not set, the monitor cannot detect drift.”
↩︎ Exam trap 1“cannot be changed after the monitor is created. You can modify only a few options using ALTER MODEL MONITOR.”
↩︎ Exam trap 3“cannot be changed after the monitor is created. You can modify only a few options using ALTER MODEL MONITOR.”
↩︎ Checkpoint“Although this parameter is optional, if it is not set, the monitor cannot detect drift.”
↩︎ Prediction“These columns are not treated as features.”
↩︎ Checkpoint - 3.
“monitoring your data for anomalies, by using Anomaly Detection functions within Snowflake Tasks or Alerts”
↩︎ What a model monitor watches, and how you create one - 4.https://docs.snowflake.com/en/user-guide/alertsOfficial docs
“Alert on a schedule: Snowflake evaluates the condition against the existing data on a scheduled basis.”
↩︎ What a model monitor watches, and how you create one - 5.
“Valid values: 'JENSEN_SHANNON' 'DIFFERENCE_OF_MEANS' 'WASSERSTEIN' 'POPULATION_STABILITY_INDEX'”
↩︎ Querying drift, stat and performance metrics for thresholds - 6.
“Each function requires the name of a model monitor and the name of a metric to be retrieved from that model.”
↩︎ Querying drift, stat and performance metrics for thresholds
Also cited
“'COUNT' and 'COUNT_NULL' are supported for all column types.”
↩︎ Checkpoint