CertSafari
    Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01)· Lessons

    Domain 4 · Lesson 14/17

    Troubleshooting Retraining Pipelines: Warehouse Capacity, Container Scaling and Data Dependencies

    Implement retraining and troubleshooting.

    10 min read
    7.34% of exam
    8 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Work through Snowflake's checklist for a retraining task that did not run
    • Diagnose canceled or skipped runs caused by warehouse sizing, timeouts and overlapping graph runs
    • Fix multi-node ML Job failures caused by compute pool size and node availability
    • Trace failures to upstream data: skipped child tasks, empty or stale streams, and suspended model monitors

    1.First pass: did the retraining task run at all?

    When retraining doesn't happen, start with TASK_HISTORY, as Snowflake's troubleshooting guide does. The table function shows whether the task ran, its scheduled and completed times, and any error code and message. A task can run successfully while the SQL in its body fails, and the history separates those two cases. To be told about failures instead of finding them later, you can create an alert on new data against the event table and filter task events with value:state = 'FAILED'.

    If the history shows no run, check the following in order. Is the task, and every task in its graph, actually RESUMED? Snowflake creates every task suspended. Resuming the root alone does not enable the children: call SYSTEM$TASK_DEPENDENTS_ENABLE on the root, or resume each child task. For a scheduled task, check the cron expression and confirm that at least one scheduled time has passed. Finally, check the owner role's privileges with SHOW GRANTS TO ROLE.

    Privileges the task owner role needs for a task to run
    ObjectPrivilegeNote
    AccountEXECUTE TASKRevoking it stops all later runs under that role
    DatabaseUSAGE
    SchemaUSAGE
    TaskOWNERSHIP
    WarehouseUSAGE

    Checkpoint 1 of 6· Check yourself

    An engineer resumes only the root task of a new retraining graph. The root runs on schedule, but training and evaluation never execute. What is the most likely cause?

    Sources12

    2.Warehouse capacity: canceled runs, timeouts and skipped graph runs

    A single task run is limited to 60 minutes by default, as a safeguard against tasks that never finish. Retraining jobs that grow with the data often reach this limit. The first fix is a larger warehouse, so the run fits within the schedule window or the one-hour limit. The second is ALTER TASK … SET USER_TASK_TIMEOUT_MS to raise the limit for that task. Neither helps if the problem is query parallelization; then the SQL itself has to be rewritten.

    Check whether a task has a custom timeout; no row means the default 3600000 mssql
    SHOW PARAMETERS LIKE 'USER_TASK_TIMEOUT_MS' IN TASK <task_name>;

    Capacity problems also appear at the level of the whole graph. Sibling tasks under one parent run in parallel, so a shared user-managed warehouse has to be sized for all of them running at once. By default only one instance of a graph runs at a time, and the next root run is scheduled only after every task has finished. If a graph takes longer than its root schedule, at least one run is skipped. The fixes are a longer interval between root runs, a larger warehouse or serverless compute, or a different overlap policy where overlapping runs can't corrupt data.

    OVERLAP_POLICY values on the root task
    OVERLAP_POLICYBehaviour
    NO_OVERLAP (default)Next root run is scheduled only after all child tasks finish
    ALLOW_CHILD_OVERLAPA new graph instance can start while children still run; root tasks never overlap
    ALLOW_ALL_OVERLAPMultiple instances of the whole graph, root included, can run concurrently

    Checkpoint 2 of 6· Match them up

    Match each symptom to the setting or fix Snowflake points to

    Tap a term, then the definition that fits it.

    Checkpoint 3 of 6· Exam question

    An engineer creates an alert that checks the model monitor's F1 metric every 15 minutes and should execute the retraining task when it falls below 0.80. Three days later retraining has never run, even though the metric has been below 0.80 for two days, and the alert has no entries in ALERT_HISTORY. What is the most likely cause?

    Sources13

    3.Container scalability: multi-node training jobs and compute pools

    Distributed retraining runs as a multi-node ML Job on a compute pool in Snowpark Container Services. The node count is set when the job is submitted, through target_instances. ML Jobs wait for that many nodes before they execute the payload. If the nodes aren't available within the timeout period, the job fails with an error. Two things cause this.

    The first is a pool ceiling that is too low. A compute pool's MAX_NODES must be greater than or equal to target_instances. A pool capped below the requested node count can never satisfy the job. The second is contention: nodes may start at different times, especially if compute pool resources are limited, and some may start seconds or minutes after others.

    Submitting a training script as a multi-node ML Job: node count is fixed at submissionpython
    from snowflake.ml.jobs import submit_file
    
    job = submit_file(
        "<script_path>",
        "MY_COMPUTE_POOL",
        stage_name="<payload_stage>",
        session=session,
        target_instances=<num_training_nodes>  # Specify the number of nodes
    )
    Quickstart compute pool sized for a two-instance jobsql
    CREATE COMPUTE POOL IF NOT EXISTS DEMO_POOL
        MIN_NODES = 1
        MAX_NODES = 2
        INSTANCE_FAMILY = CPU_X64_S;

    When contention, not the ceiling, is the problem, set the optional min_instances parameter. The job then starts as soon as the minimum number of nodes is up, even if that is fewer than target_instances. This shortens waits on a busy pool and lets a retraining pipeline adapt to the resources it gets. The training code has to be written for distributed execution, using Snowflake's Distributed Modeling Classes or Ray, so that it runs correctly on however many nodes it receives.

    Checkpoint 4 of 6· Check yourself

    A retraining graph submits an ML Job with target_instances=4 to a compute pool created with MAX_NODES = 2. What happens?

    Sources45

    4.Data dependencies: skipped children, empty or stale streams, suspended monitors

    Some retraining failures start upstream of training. In a task graph, a child runs only after its parents finish. If a parent fails to run to completion, any child tasks are skipped, so a skipped training task often points to a failed feature-preparation task. Check the predecessor before the task itself.

    For stream-triggered retraining, the dependency is the stream. If the task has a SYSTEM$STREAM_HAS_DATA condition, check that the stream actually held change records when the task was last scheduled; an AT | BEFORE clause queries the stream's history. Streams can also go stale: the task must consume stream data before data retention expires. For a stream on a directory table, the directory table has to be refreshed (by auto-refresh or ALTER STAGE … REFRESH) before a triggered task can detect new files.

    The drift signal itself depends on data. A model monitor automatically suspends refreshes after five consecutive refresh failures related to its source tables. A suspended monitor stops updating, so any drift-based policy that reads it is working from old data. DESCRIBE MODEL MONITOR shows what happened. In aggregation_status, one or more values are SUSPENDED. aggregation_last_error holds the specific SQL error. The source and baseline JSON objects carry a status that can be DELETED or MASKED, meaning the table was dropped or the current role can't see it.

    Checkpoint 5 of 6· Put it in order

    Drift readings stopped changing a week ago. Put the recovery steps in order.

    1. 1.Run ALTER MODEL MONITOR … RESUME
    2. 2.Fix the root cause in the source tables
    3. 3.Run DESCRIBE MODEL MONITOR and find SUSPENDED values in aggregation_status
    4. 4.Read aggregation_last_error to get the SQL error that caused the suspension

    Checkpoint 6 of 6· Exam question

    New labeled outcome rows are loaded into `labels_raw` by a pipeline at irregular times. A data scientist wants the retraining task graph to start only when new rows exist, and does not want to define a polling schedule on the root task. Which configuration is correct?

    Sources1678

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A retraining task that keeps timing out can always be fixed with a bigger warehouse or a higher USER_TASK_TIMEOUT_MS.Why is that wrong?

      Both fixes fail when the bottleneck is query parallelization. In that case the task's SQL has to be rewritten.

      Covered in Warehouse capacity: canceled runs, timeouts and skipped graph runs

    2. 2.If a multi-node ML Job can't get all its nodes, it starts with whatever nodes are available.Why is that wrong?

      By default the job waits for target_instances and fails at the timeout. It starts early only when min_instances is set.

      Covered in Container scalability: multi-node training jobs and compute pools

    3. 3.A model monitor that hits source-table errors keeps retrying, so drift numbers catch up once the data is fixed.Why is that wrong?

      After five consecutive source-related refresh failures, the monitor suspends refreshes. It stays suspended until you fix the cause and run ALTER MODEL MONITOR … RESUME.

      Covered in Data dependencies: skipped children, empty or stale streams, suspended monitors

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Query the TASK_HISTORY table function to verify the task did not run.”
      ↩︎ First pass: did the retraining task run at all?
      “Snowflake creates all tasks in the SUSPENDED state.”
      ↩︎ First pass: did the retraining task run at all?
      “Revoking the EXECUTE TASK privilege on a role prevents all subsequent task runs from starting under that role.”
      ↩︎ First pass: did the retraining task run at all?
      “If the task was canceled or exceeded the window scheduled for the task, the cause is often an undersized warehouse.”
      ↩︎ Warehouse capacity: canceled runs, timeouts and skipped graph runs
      “If the statement returns no record, the task currently has the default 3600000 millisecond (60 minute) timeout.”
      ↩︎ Warehouse capacity: canceled runs, timeouts and skipped graph runs
      “If a parent task failed to run to completion, any child tasks are skipped.”
      ↩︎ Data dependencies: skipped children, empty or stale streams, suspended monitors
      “verify that the specified stream contained change data capture (CDC) records when the task was last scheduled to run.”
      ↩︎ Data dependencies: skipped children, empty or stale streams, suspended monitors
      “neither increasing the warehouse size nor increasing the timeout limit might help if there are query parallelization issues.”
      ↩︎ Exam trap 1
      “To recursively enable all dependent tasks tied to a root task, query the SYSTEM$TASK_DEPENDENTS_ENABLE function rather than enabling each task individually.”
      ↩︎ Checkpoint
    2. 2.
      “To return events where the task execution failed, use value:state = 'FAILED'.”
      ↩︎ First pass: did the retraining task run at all?
    3. 3.
      “at least one run of the task graph is skipped.”
      ↩︎ Warehouse capacity: canceled runs, timeouts and skipped graph runs
      “If feasible, increase the scheduling time between runs of the root task.”
      ↩︎ Warehouse capacity: canceled runs, timeouts and skipped graph runs
      “When tasks run in parallel on the same user-managed warehouse, the compute resources must be sized to handle the concurrent task runs.”
      ↩︎ Checkpoint
    4. 4.
      “ML Jobs automatically wait for the specified target_instances to be available before executing your payload.”
      ↩︎ Container scalability: multi-node training jobs and compute pools
      “Nodes might not all start simultaneously, especially if compute pool resources are limited”
      ↩︎ Container scalability: multi-node training jobs and compute pools
      “the job payload is executed as soon as the minimum number of nodes becomes available, even if that number is smaller than target_instances.”
      ↩︎ Container scalability: multi-node training jobs and compute pools
      “The job fails with an error if the expected nodes aren’t available within the timeout period.”
      ↩︎ Exam trap 2
      “You must set MAX_NODES to be greater than or equal to the number of target instances that you’re using to run your training job.”
      ↩︎ Checkpoint
    5. 5.
      “Note: MAX_NODES should be at least equal to target_instances (2 in this example).”
      ↩︎ Container scalability: multi-node training jobs and compute pools
    6. 6.
      “Task instructions must consume stream data before data retention expires; otherwise, the stream becomes stale.”
      ↩︎ Data dependencies: skipped children, empty or stale streams, suspended monitors
      “A directory table must be refreshed before a triggered task can detect the changes.”
      ↩︎ Data dependencies: skipped children, empty or stale streams, suspended monitors
    7. 7.
      “aggregation_last_error: The value in this column is a JSON object that contains the specific SQL error that caused the suspension.”
      ↩︎ Data dependencies: skipped children, empty or stale streams, suspended monitors
      “Model monitors automatically suspend refreshes when they encounter five consecutive refresh failures related to the source tables.”
      ↩︎ Exam trap 3
      “After resolving the root cause of the refresh failure, resume the monitor by issuing ALTER MODEL MONITOR … RESUME.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 26 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.