What you will be able to do
- Build a task graph whose root, child and finalizer tasks chain validation, training and deployment
- Pass values between tasks and start a retraining graph from a stream instead of a polling schedule
- Rerun a failed graph from the failed task and control retries, suspension and overlap
- Pick Snowflake CLI over legacy SnowSQL to deploy and run pipeline SQL from scripts and CI
1.Wiring stages into a task graph
Inside Snowflake, a pipeline such as validate → train → evaluate → deploy usually runs as a task graph, also called a DAG. It has a root task, dependent child tasks, and optionally a finalizer. Dependencies run from start to finish, with no loops. You can define tasks in SQL, Python, Java, Scala, JavaScript or Snowflake Scripting, and you can view and manage the graph in Snowsight.
You create the root task first, then create each child with CREATE TASK … AFTER to name its parents. The root task decides when the graph runs. Children run in the order the graph defines. When several children share a parent, they run in parallel. When a task has several parents, it waits for all of them to complete successfully.
CREATE TASK task_root
SCHEDULE = '1 MINUTE'
AS SELECT 1;
CREATE TASK task_a
AFTER task_root
AS SELECT 1;
CREATE TASK task_b
AFTER task_root
AS SELECT 1;
CREATE TASK task_c
AFTER task_a, task_b
AS SELECT 1;Limits and rules: a graph can hold at most 1000 tasks, and a single task can have at most 100 parents and 100 children. Every task in the graph must have the same owner and live in the same database and schema. If parallel tasks share a user-managed warehouse, size the warehouse for the concurrent runs.
A finalizer, created with CREATE TASK … FINALIZE = <root>, runs after all other tasks have completed or failed. Use it to clean up intermediate data or send success and failure notifications. Each root has at most one finalizer. A finalizer cannot have children, and it does not start if the root task was skipped.
To test a graph without enabling its schedule, resume the child tasks you want included (including the finalizer), then run EXECUTE TASK on the root. To start the graph for real, resume the children and then the root. Alternatively, call SYSTEM$TASK_DEPENDENTS_ENABLE on the root to resume every task at once.
Checkpoint 1 of 7· Put it in order
Put these steps in order to build and start a scheduled task graph manually
- 1.Create each child task with CREATE TASK … AFTER its parent tasks
- 2.Create the root task with CREATE TASK and a SCHEDULE
- 3.Resume the root task with ALTER TASK … RESUME
- 4.Resume each child task, including the finalizer, with ALTER TASK … RESUME
Parents must exist before a child can name them in AFTER. Children are resumed before the root, because resuming the root is what starts scheduled runs.
“Create a root task using CREATE TASK, then create child tasks using CREATE TASK .. AFTER to select the parent tasks.”Source: docs.snowflake.com
Checkpoint 2 of 7· Check yourself
TRAIN_MODEL is defined with AFTER VALIDATE_SCHEMA, VALIDATE_FRESHNESS. Both validation tasks are resumed. When does TRAIN_MODEL start?
A task with several parents waits for all of them, so the training step acts as a gate on both data checks.
“When a task has multiple parents, the task waits for all preceding tasks to successfully complete before starting.”Source: docs.snowflake.com
Sources1
2.Passing results between tasks and triggering on new data
In a real pipeline, a downstream task often needs something an upstream task produced, such as a new model version name or an evaluation score. Task graphs support this with return values. A task calls SYSTEM$SET_RETURN_VALUE('<string>'). Any task that names it as a predecessor in AFTER can then read the value with SYSTEM$GET_PREDECESSOR_RETURN_VALUE. The value is a string of at most 10 kB in UTF-8.
Snowflake's end-to-end task graph guide uses this pattern. Its graph prepares data, trains on distributed compute, evaluates the model against quality thresholds, and only then promotes it, finishing with notifications and cleanup. Promotion is conditional: the deploy step runs logic based on what evaluation reported.
Checkpoint 3 of 7· Check yourself
In one graph run, TRAIN_MODEL creates a new model version name, and DEPLOY_MODEL (defined AFTER TRAIN_MODEL) has to use it. What is the supported mechanism?
Return values are the built-in way to pass data from a parent task to the child tasks that list it in AFTER.
“can retrieve the return value set by the predecessor task using SYSTEM$GET_PREDECESSOR_RETURN_VALUE.”Source: docs.snowflake.com
The root task decides *when* the graph runs: on a recurring schedule, or when an event fires. For retraining on newly arrived rows, a schedule that polls the table wastes runs whenever nothing has changed. A triggered task runs when a stream changes. You define the target stream in the WHEN clause and leave out the SCHEDULE parameter. A triggered task uses no compute until the event fires, and it also lowers latency because new data is processed right away. Triggered tasks work with streams on tables, views, dynamic tables, Iceberg tables, data shares and directory tables. They do not work with hybrid tables or with streams on external tables.
Checkpoint 4 of 7· Exam question
Which statements about a finalizer task in a Snowflake task graph are true? (Select two.)(Select 2)
Correct answers: A, E — Each root task can have only one finalizer task, and that finalizer task cannot itself have any child tasks below it in the graph.; A finalizer runs after all other tasks in the graph run have finished, succeeded or failed, so it suits cleanup such as dropping temporary tables.
- A. Correct. Both restrictions are part of the finalizer definition for a task graph.
- B. Incorrect. If the root task is skipped, the finalizer does not start for that run.
- C. Incorrect. A finalizer also runs after failures, so promotion placed there could publish a model from a broken run.
- D. Incorrect. A finalizer is attached to the root task and runs as part of each graph run, not on a separate schedule.
- E. Correct. This is the purpose of a finalizer: guaranteed post-run work regardless of the outcome of the other tasks.
3.Failures, reruns and overlapping runs
By default, if any child task fails, the whole graph run counts as failed. After a fix, you don't need to start again from the root. EXECUTE TASK … RETRY LAST reruns the latest graph run starting from the last failed task, and its children continue as their predecessors complete. To retry an older run, use EXECUTE TASK … RETRY GRAPH RUN GROUP with that run's GRAPH_RUN_GROUP_ID. To find the failure, call the TASK_DEPENDENTS table function on the root to list the graph's tasks, or open the graph in Snowsight.
Two parameters on the root task control automatic behaviour. TASK_AUTO_RETRY_ATTEMPTS retries the graph immediately when a child fails, instead of waiting for the next scheduled run. SUSPEND_TASK_AFTER_NUM_FAILURES suspends the graph after a run of consecutive failures. The default is 10.
CREATE OR REPLACE TASK task_root
SCHEDULE = '1 MINUTE'
TASK_AUTO_RETRY_ATTEMPTS = 2 -- Failed task graph retries up to 2 times
SUSPEND_TASK_AFTER_NUM_FAILURES = 3 -- Task graph suspends after 3 consecutive failures
AS SELECT 1;To change any task in a scheduled graph, first suspend the root. A run already in progress finishes, and future scheduled runs are cancelled. To skip one step, suspend that child task. The graph then continues as if the child had succeeded.
By default, only one instance of a graph runs at a time. If a run takes longer than the root's schedule interval, at least one scheduled run is skipped. The root task's OVERLAP_POLICY changes this behaviour:
| OVERLAP_POLICY | Behaviour |
|---|---|
| NO_OVERLAP (default) | Next root run is scheduled only after all child tasks finish |
| ALLOW_CHILD_OVERLAP | A new graph instance can start while children are still running; root tasks never overlap |
| ALLOW_ALL_OVERLAP | Multiple instances of the entire graph, including the root, can run concurrently |
Checkpoint 5 of 7· Check yourself
In last night's run, VALIDATE_DATA succeeded and TRAIN_MODEL failed. The bug is fixed. How do you rerun from the failed task without repeating validation?
RETRY LAST restarts the latest graph run from the last failed task. A plain EXECUTE TASK on the root starts a whole new run, and the auto-retry parameter only affects future failures.
“To retry the latest graph run, use EXECUTE TASK … RETRY LAST to attempt to run the task graph from the last failed task.”Source: docs.snowflake.com
Checkpoint 6 of 7· Exam question
A team defined a graph of six tasks with CREATE TASK ... AFTER statements. They want every task in the graph enabled with one call after deployment. Which approach does that?
Correct answer: B — Run SELECT SYSTEM$TASK_DEPENDENTS_ENABLE on the root task's name, which resumes the root along with all of its dependent tasks.
- A. Incorrect. Resuming the root does not resume children; suspended children are skipped when the graph runs.
- B. Correct. This system function enables the root and every dependent task in one call.
- C. Incorrect. The OPERATE privilege controls who may resume tasks; it never resumes any task by itself.
- D. Incorrect. EXECUTE TASK triggers a run; it does not change the suspended or started state of child tasks.
Sources1
4.Driving the pipeline from Snowflake CLI, SnowSQL and SDKs
Pipeline definitions such as task DDL, stage code and SQL scripts are deployed and run from developer tooling, not typed into a worksheet. Snowflake CLI (snow) is the preferred command-line client for new installs and new work. It is open source and built for developer workloads as well as SQL, covering Snowpark, Snowpark Container Services, notebooks, stages, Git repositories and SQL execution. SnowSQL is the legacy client. It receives no new features, and SnowSQL 1.5.x is supported through April 16, 2028. Teams still on SnowSQL have a documented migration path to Snowflake CLI.
uv tool install snowflake-cli
snow --versionIn an automated deployment, the CLI does the deploy step. snow sql -f <file> runs standalone SQL scripts, such as the file that creates your task graph. snow snowpark deploy deploys Snowpark applications. snow git execute runs SQL files from a Snowflake Git integration. The language SDKs cover what goes on inside the steps: Snowpark Python and snowflake-ml-python build the session and submit ML Jobs. Because ML Jobs can be submitted from externally hosted orchestrators, a team can keep its own scheduler and still run compute in Snowflake.
Checkpoint 7 of 7· Check yourself
A team is starting a new project to script deployment of its task graph SQL. Which command-line client should it use?
SnowSQL is legacy and stays supported for a limited time, but new features and enhancements only go into Snowflake CLI.
“SnowSQL is a legacy command-line client. Snowflake will only add new features and enhancements to Snowflake CLI.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Suspending a child task makes the task graph stop at that point.Why is that wrong?
A suspended child is skipped. The graph keeps running as though that child had succeeded.
Covered in Failures, reruns and overlapping runs
2.A stream-triggered retraining root task also needs a short SCHEDULE so it can check the stream.Why is that wrong?
A triggered task names its stream in the WHEN clause and must not include SCHEDULE. It uses no compute until the event fires.
Covered in Passing results between tasks and triggering on new data
3.Tasks in one graph can live in different schemas, for example validation tasks in a DATA schema and training tasks in an ML schema.Why is that wrong?
Every task in a graph must have the same owner and be stored in the same database and schema.
Covered in Wiring stages into a task graph
Practise it for real
Build the four-task sample graph with a finalizer, then run it once without enabling its schedule
1.Run the CREATE TASK statements for task_root, task_a, task_b and task_c from the task-graph sample.
Why: These statements create the root, the two parallel children and the join task.
You should see: Four tasks exist. task_a and task_b list task_root as their parent, and task_c lists both of them.
2.Run CREATE TASK task_finalizer FINALIZE = task_root AS SELECT 1;
Why: A finalizer is where a pipeline does its cleanup and sends notifications.
You should see: The finalizer is attached to task_root. It can have no child tasks.
3.Call the TASK_DEPENDENTS table function with task_root as input.
Why: Lists every task in the graph, so you can check the wiring.
You should see: All tasks in the graph are returned.
4.Run ALTER TASK … RESUME on task_a, task_b, task_c and task_finalizer, then EXECUTE TASK task_root.
Why: EXECUTE TASK on the root runs a single instance of the graph, which is how you test before production.
You should see: The resumed children run in graph order: task_a and task_b in parallel, then task_c, then the finalizer.
Stuck? Get a nudge
If a child does not run, check whether you resumed it. EXECUTE TASK only runs child tasks that are resumed.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“When multiple child tasks have the same parent, the child tasks run in parallel.”
↩︎ Wiring stages into a task graph“To run a single instance of a task graph, use EXECUTE TASK on the root task.”
↩︎ Wiring stages into a task graph“A finalizer task cannot have any child tasks.”
↩︎ Wiring stages into a task graph“By default, a task graph is suspended after 10 consecutive failures.”
↩︎ Failures, reruns and overlapping runs“To modify a task in a scheduled task graph, suspend the root task using ALTER TASK … SUSPEND.”
↩︎ Failures, reruns and overlapping runs“OVERLAP_POLICY = NO_OVERLAP (default): Executes tasks serially with no parallelism.”
↩︎ Failures, reruns and overlapping runs“When you suspend a child task, the task graph continues to run as though the child task had succeeded.”
↩︎ Exam trap 1“All tasks in a task graph must have the same task owner and be stored in the same database and schema.”
↩︎ Exam trap 3“Create a root task using CREATE TASK, then create child tasks using CREATE TASK .. AFTER to select the parent tasks.”
↩︎ Checkpoint“When a task has multiple parents, the task waits for all preceding tasks to successfully complete before starting.”
↩︎ Checkpoint“To retry the latest graph run, use EXECUTE TASK … RETRY LAST to attempt to run the task graph from the last failed task.”
↩︎ Checkpoint - 2.
“In a task graph, a task can call this function to set a return value.”
↩︎ Passing results between tasks and triggering on new data“The string size must be <= 10 kB (when encoded in UTF8).”
↩︎ Passing results between tasks and triggering on new data“can retrieve the return value set by the predecessor task using SYSTEM$GET_PREDECESSOR_RETURN_VALUE.”
↩︎ Checkpoint - 3.https://www.snowflake.com/en/developers/guides/e2e-task-graphSecondary source
“Conditionally promotes high-quality models to production in Model Registry”
↩︎ Passing results between tasks and triggering on new data - 4.
“Triggered tasks don’t use compute resources until the event is triggered.”
↩︎ Passing results between tasks and triggering on new data“Define the target stream using the WHEN clause. (Do not include the SCHEDULE parameter.)”
↩︎ Exam trap 2 - 5.
“Snowflake CLI is the preferred command-line client for new installs and new work.”
↩︎ Driving the pipeline from Snowflake CLI, SnowSQL and SDKs“Snowflake supports SnowSQL 1.5.x through April 16, 2028.”
↩︎ Driving the pipeline from Snowflake CLI, SnowSQL and SDKs“SnowSQL is a legacy command-line client. Snowflake will only add new features and enhancements to Snowflake CLI.”
↩︎ Checkpoint - 6.
“snow git execute to run SQL files from a Snowflake Git integration”
↩︎ Driving the pipeline from Snowflake CLI, SnowSQL and SDKs