What you will be able to do
- Read drift signals from a model version monitor with MODEL_MONITOR_DRIFT_METRIC and know what the monitor needs before it can compute them
- Express a retraining policy as a Snowflake alert: a condition, an action and a schedule
- Choose between an alert on a schedule, an alert on new data and a stream-triggered task to start retraining
- Set retry and auto-suspend behaviour on a retraining task graph
- Troubleshoot a failed retraining pipeline by separating warehouse-capacity, container-scaling and data-dependency causes and applying the matching fix
Key concept
Alert-driven retraining policy — A retraining policy in Snowflake is a schema-level alert. Its condition queries a monitoring signal such as a drift metric, its action starts the retraining work, and its schedule says how often the condition is checked. The policy lives in the alert definition, so nobody has to watch a dashboard and start retraining by hand.
1.Where the retraining signal comes from: model version monitors
An automated retraining policy needs a number it can test. In Snowflake that number comes from ML Observability. A model monitor lets you track the quality of production models you have deployed via the Snowflake Model Registry across multiple dimensions, such as performance, drift, and volume. Each monitor watches exactly one model version. It reads a table of stored inference data (an ID, a timestamp, features, predictions and ground-truth labels), refreshes its logs from that table, and aggregates them over a time window whose minimum is one day.
That one-day minimum affects policy design. A drift signal cannot be finer than the monitor's aggregation window, so a policy that checks drift every minute just reads the same daily figure over and over.
Drift is read with the MODEL_MONITOR_DRIFT_METRIC table function. You give it the monitor, a metric name and a column. The column can be a feature, the prediction or the actual for model version monitors. Optional arguments set the granularity, the time range and a single segment filter. Each output row has an EVENT_TIMESTAMP and a METRIC_VALUE, plus the counts of records used and excluded.
| drift_metric_name value | Argument it fills |
|---|---|
| 'JENSEN_SHANNON' | drift_metric_name |
| 'DIFFERENCE_OF_MEANS' | drift_metric_name (only metric with a CI_VALUE, binary classification only) |
| 'WASSERSTEIN' | drift_metric_name |
| 'POPULATION_STABILITY_INDEX' | drift_metric_name |
SELECT * FROM TABLE(MODEL_MONITOR_DRIFT_METRIC(
'MY_MONITOR', 'JENSEN_SHANNON', 'MODEL_PREDICTION', '1 DAY', DATEADD('DAY', -30, CURRENT_DATE()), CURRENT_DATE())
)Drift is a comparison, so the monitor needs something to compare against: an optional baseline table set on the monitor. Without a baseline, the function returns an error, not a value. Other ways to get an error are asking for a numerical metric on a non-numeric feature and asking for a metric the monitor doesn't have. The monitor also tracks performance as its own dimension, but these sources document only the drift function. A performance-degradation policy uses the same alert pattern described below, with the monitor's performance data as its condition.
Checkpoint 1 of 6· Check yourself
A team creates a model version monitor without a baseline table and then writes a drift-based retraining condition with MODEL_MONITOR_DRIFT_METRIC. What happens when the condition runs?
Drift metrics need a baseline set on the model version monitor. Without one, the call errors. There is no implicit or fallback baseline.
“The model version monitor must have a baseline set for the drift metric to be computed.”Source: docs.snowflake.com
2.Writing the policy as an alert
An alert has three parts: a condition that triggers it, an action to perform when the condition is met, and a schedule for how often the condition is checked. For retraining, the condition is the policy itself, for example "Jensen-Shannon drift on the prediction column above an agreed threshold in the latest window". The action is what happens next. SUSPEND_ALERT_AFTER_NUM_FAILURES keeps a broken policy from failing quietly forever.
CREATE [ OR REPLACE ] ALERT [ IF NOT EXISTS ] <name>
[ [ WITH ] TAG ( <tag_name> = '<tag_value>' [ , <tag_name> = '<tag_value>' , ... ] ) ]
[ SCHEDULE = '{ <num> MINUTE | USING CRON <expr> <time_zone> }' ]
[ WAREHOUSE = <warehouse_name> ]
[ COMMENT = '<string_literal>' ]
[ CONFIG = '<configuration_string>' ]
[ RUNBOOK = '<string_literal>' ]
[ SUSPEND_ALERT_AFTER_NUM_FAILURES = <number> ]
IF( EXISTS(
<condition>
))
THEN
<action>There are two kinds of alert, and the kind decides what the condition can contain. An alert on a schedule evaluates its condition against all existing data, every n minutes or on a cron expression. An alert on new data evaluates only rows newly inserted into one table or view. That table or view must have change tracking enabled, and the condition cannot use CTEs, DML, stored-procedure calls or joins. Because the new-data FROM clause is limited to one regular table, view or event table, a condition that calls the MODEL_MONITOR_DRIFT_METRIC table function belongs in an alert on a schedule. An alert on new data suits a condition such as "new rows landed in a labelled-feedback table".
The action can read what the condition found. It calls GET_CONDITION_QUERY_UUID and passes the result to RESULT_SCAN, so it can log the drift value that triggered it. The documented actions are SQL, such as calling SYSTEM$SEND_EMAIL or running a BEGIN … END block. A retraining policy can therefore send its action to the retraining task graph. Running EXECUTE TASK on the graph's root task starts a single run of every resumed child task.
| Property | Alert on a schedule | Alert on new data |
|---|---|---|
| What is evaluated | All existing data, each run | Only newly inserted rows |
| When it runs | Every n MINUTE or USING CRON | When new rows are inserted |
| Condition limits | None of the new-data restrictions | One table/view/event table; change tracking; no CTEs, DML, procedure calls or joins |
| Manual test with EXECUTE ALERT | Allowed | Not allowed |
Checkpoint 2 of 6· Check yourself
A retraining policy should fire when daily Jensen-Shannon drift from MODEL_MONITOR_DRIFT_METRIC passes a threshold. Which alert design fits the documented restrictions?
An alert on new data allows only one regular table, view or event table in FROM, and it cannot use joins or EXECUTE ALERT. A table-function drift check therefore belongs in a scheduled alert.
“In the SELECT statement, the FROM clause can specify only one regular table, view, or event table.”Source: docs.snowflake.com
Checkpoint 3 of 6· Exam question
A model monitor on a fraud-scoring model in Snowflake records drift metrics hourly. The team wants retraining to start automatically when the Jensen-Shannon drift of the `txn_amount` feature exceeds a threshold, with no always-on external service. Which design meets this requirement?
Correct answer: A — Create an alert whose condition queries the monitor's drift metric function for values above the threshold, and whose action runs EXECUTE TASK on the retraining root task.
- A. An alert evaluates its condition query on a schedule and runs a SQL action only when rows are returned, so checking the drift metric and executing the task gives a threshold-driven trigger entirely inside Snowflake.
- B. A stream-triggered task fires whenever new rows arrive, which is data arrival rather than measured drift, so it would retrain on every batch and ignore the threshold.
- C. A fixed nightly schedule retrains whether or not drift occurred and cannot respond to drift between runs, so it is a calendar policy rather than a drift-triggered one.
- D. Dynamic tables only materialize query results and cannot invoke stored procedures on refresh, so this cannot start training and it also ignores any drift threshold.
3.What the policy starts: triggered tasks and a self-healing retraining graph
An alert isn't the only trigger. If the policy is "retrain when new labelled data arrives", a triggered task can watch a stream on the training-data table. Triggered tasks don't use compute resources until the event is triggered, and they work on streams over tables, views, dynamic tables, Iceberg tables, data shares and directory tables. They don't work with hybrid tables or with streams on external tables. You define one with a WHEN clause and no SCHEDULE. A serverless triggered task also needs TARGET_COMPLETION_INTERVAL, which Snowflake uses to size compute so the run finishes within that interval. A triggered task runs at most every 30 seconds by default, and only one instance runs at a time.
Checkpoint 4 of 6· Fill the gap
Which keyword makes this serverless task run whenever the stream has data, instead of on a timer?
CREATE TASK my_triggered_task
TARGET_COMPLETION_INTERVAL='15 MINUTES'
? SYSTEM$STREAM_HAS_DATA('my_order_stream')
AS
INSERT INTO customer_activity
SELECT customer_id, order_total, order_date, 'order'
FROM my_order_stream;Triggered tasks name their stream in the WHEN clause and must not include SCHEDULE. AFTER links a child to a parent, and IF belongs to alert syntax.
Source: docs.snowflake.comWhatever starts it, retraining usually runs as a task graph: prepare data, train, evaluate against thresholds, and promote conditionally. Snowflake's end-to-end example graph "conditionally promotes high-quality models to production in Model Registry". The quality gate between tasks can be built in the task body from runtime values, graph-level configuration and the return values of parent tasks.
A policy also needs to say what happens when retraining fails. Two parameters on the root task set this: how many times a failed graph retries, and after how many consecutive failures the whole graph suspends.
CREATE OR REPLACE TASK task_root
SCHEDULE = '1 MINUTE'
TASK_AUTO_RETRY_ATTEMPTS = 2 -- Failed task graph retries up to 2 times
SUSPEND_TASK_AFTER_NUM_FAILURES = 3 -- Task graph suspends after 3 consecutive failures
AS SELECT 1;Checkpoint 5 of 6· Put it in order
A nightly scheduled retraining task should instead run whenever a stream on new labelled data has rows. Put the conversion steps in order.
- 1.ALTER TASK … MODIFY WHEN SYSTEM$STREAM_HAS_DATA(...)
- 2.ALTER TASK … RESUME
- 3.ALTER TASK … UNSET SCHEDULE
- 4.ALTER TASK … SUSPEND
You suspend the task, remove its schedule, add the WHEN clause for the stream, and then resume it. A triggered task cannot keep a SCHEDULE.
“Use ALTER TASK to update the task. Unset the SCHEDULE parameter, and then add the WHEN clause to define the target stream.”Source: docs.snowflake.com
4.Troubleshooting retraining pipeline failures: warehouse capacity, container scaling and data dependencies
When a retraining pipeline fails, first work out which of three families the failure belongs to. Start with TASK_HISTORY: check the scheduled and completed times and any error code and message, because a task can run successfully even though the SQL inside it failed.
Warehouse capacity. A single task run has a 60 minute default limit. If a task was canceled or overran its schedule window, the cause is often an undersized warehouse, so review the warehouse size and consider increasing it. Alternatively, raise the limit with ALTER TASK … SET USER_TASK_TIMEOUT_MS. A bigger warehouse or longer timeout might not help if there are query parallelization issues; then rewrite the SQL. In a graph, child tasks that share a parent run in parallel, and when they run on the same user-managed warehouse it must be sized for the concurrent runs. If the whole graph takes longer than the root task's schedule, at least one run is skipped, so lengthen the schedule interval, enlarge the warehouse or use serverless compute.
Container scalability. Training that runs as a multi-node ML Job on a compute pool needs MAX_NODES at least equal to the target instances. ML Jobs wait for the target_instances nodes, and the job fails with an error if they are not available within the timeout. Setting min_instances lets the job start with fewer nodes when compute pool resources are limited. For a pool that cannot be provisioned because of an insufficient capacity error, BACKUP_INSTANCE_FAMILIES lists fallback instance families. For services, autoscaling only happens when MAX_INSTANCES is greater than MIN_INSTANCES, and a service with no CPU and memory requirements in its specification does not autoscale.
Data dependencies. If a parent task in a graph fails to run to completion, its child tasks are skipped, so look at the predecessor first. For a task with a SYSTEM$STREAM_HAS_DATA condition, verify the stream held change data capture records when the task was last scheduled. Model monitors have their own dependency on source tables: they suspend refreshes after five consecutive refresh failures related to the source tables. DESCRIBE MODEL MONITOR shows aggregation_status and aggregation_last_error, and after fixing the root cause you run ALTER MODEL MONITOR … RESUME. Once the cause is fixed, EXECUTE TASK … RETRY LAST re-runs a failed graph from the failed tasks.
SHOW PARAMETERS LIKE 'USER_TASK_TIMEOUT_MS' IN TASK <task_name>;Checkpoint 6 of 6· Check yourself
A nightly retraining task is canceled after running for an hour on a small warehouse. Which diagnosis and first remedy match the documentation?
A cancelled run that exceeds the window or the one-hour default points to warehouse capacity. Size up the warehouse or raise the task timeout; the other options describe different failure families.
“the cause is often an undersized warehouse”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A model version monitor can always return drift metrics once it has inference logs.Why is that wrong?
Drift needs a baseline table set on the monitor. Without one, MODEL_MONITOR_DRIFT_METRIC returns an error, so a drift-based policy on that monitor can never fire.
Covered in Where the retraining signal comes from: model version monitors
2.Any alert can be tested on demand with EXECUTE ALERT.Why is that wrong?
EXECUTE ALERT works for scheduled alerts only. An alert on new data runs only when new rows are inserted.
Covered in Writing the policy as an alert
3.A compute pool for a multi-node training job can be sized with any MAX_NODES, because the job simply uses whatever nodes exist.Why is that wrong?
MAX_NODES must be at least the number of target instances; requesting more nodes than the pool provides can make the job fail or behave unpredictably.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/model-observabilityOfficial docs
“track the quality of production models you have deployed via the Snowflake Model Registry across multiple dimensions, such as performance, drift, and volume”
↩︎ Where the retraining signal comes from: model version monitors“The minimum time granularity at which data is stored (aggregation window), currently 1 day minimum.”
↩︎ Where the retraining signal comes from: model version monitors“Model monitors automatically suspend refreshes when they encounter five consecutive refresh failures related to the source tables.”
↩︎ Troubleshooting retraining pipeline failures: warehouse capacity, container scaling and data dependencies - 2.
“Request a numerical drift metric for a non-numeric feature.”
↩︎ Where the retraining signal comes from: model version monitors“The model version monitor must have a baseline set for the drift metric to be computed.”
↩︎ Exam trap 1“The model version monitor must have a baseline set for the drift metric to be computed.”
↩︎ Checkpoint - 3.https://docs.snowflake.com/en/user-guide/alertsOfficial docs
“Alert on a schedule: Snowflake evaluates the condition against the existing data on a scheduled basis.”
↩︎ Writing the policy as an alert“Call the GET_CONDITION_QUERY_UUID function to get the query ID for the SQL statement for the condition.”
↩︎ Writing the policy as an alert“alerts are automatically suspended after the specified number of consecutive alert runs that either fail or time out”
↩︎ Writing the policy as an alert“A Snowflake alert is a schema-level object that specifies:”
↩︎ Key concept“You cannot use the EXECUTE ALERT command to execute an alert on new data.”
↩︎ Exam trap 2“In the SELECT statement, the FROM clause can specify only one regular table, view, or event table.”
↩︎ Checkpoint - 4.
“To run a single instance of a task graph, use EXECUTE TASK on the root task.”
↩︎ Writing the policy as an alert“specifying logic-based operations in the task body using runtime values, graph level configuration, and return values of parent tasks”
↩︎ What the policy starts: triggered tasks and a self-healing retraining graph“whenever a child task fails, the task graph immediately retries twice before the entire task graph is considered failed”
↩︎ What the policy starts: triggered tasks and a self-healing retraining graph“When tasks run in parallel on the same user-managed warehouse, the compute resources must be sized to handle the concurrent task runs.”
↩︎ Troubleshooting retraining pipeline failures: warehouse capacity, container scaling and data dependencies - 5.
“Triggered tasks don’t use compute resources until the event is triggered.”
↩︎ What the policy starts: triggered tasks and a self-healing retraining graph“Snowflake ensures only one instance of a task runs at a time.”
↩︎ What the policy starts: triggered tasks and a self-healing retraining graph“Use ALTER TASK to update the task. Unset the SCHEDULE parameter, and then add the WHEN clause to define the target stream.”
↩︎ Checkpoint - 6.https://www.snowflake.com/en/developers/guides/e2e-task-graphSecondary source
“Conditionally promotes high-quality models to production in Model Registry”
↩︎ What the policy starts: triggered tasks and a self-healing retraining graph - 7.https://docs.snowflake.com/en/user-guide/tasks-tsOfficial docs
“the cause is often an undersized warehouse”
↩︎ Troubleshooting retraining pipeline failures: warehouse capacity, container scaling and data dependencies“If a parent task failed to run to completion, any child tasks are skipped.”
↩︎ Troubleshooting retraining pipeline failures: warehouse capacity, container scaling and data dependencies - 8.
“The job fails with an error if the expected nodes aren’t available within the timeout period.”
↩︎ Troubleshooting retraining pipeline failures: warehouse capacity, container scaling and data dependencies“You must set MAX_NODES to be greater than or equal to the number of target instances that you’re using to run your training job.”
↩︎ Exam trap 3 - 9.https://docs.snowflake.com/en/developer-guide/snowpark-container-services/scaling-servicesOfficial docs
“Autoscaling occurs when the specified MAX_INSTANCES is greater than MIN_INSTANCES.”
↩︎ Troubleshooting retraining pipeline failures: warehouse capacity, container scaling and data dependencies
Also cited
“Newly created or cloned alerts are suspended upon creation.”
↩︎ Prediction