What you will be able to do
- Create a compute pool for ML work and explain what MIN_NODES, MAX_NODES and INSTANCE_FAMILY control
- Use AUTO_SUSPEND_SECS and AUTO_RESUME to control idle compute pool cost
- Grant the right compute pool privilege (USAGE, OPERATE, MODIFY, MONITOR) for a given task
- Decide whether an ML workload belongs on a Snowpark-optimized warehouse or on Container Runtime
- Deploy a custom container image as a Custom Runtime Environment and reference it from an ML Job
Key concept
Container Runtime on compute pools — Snowflake ML workloads such as training, tuning and batch inference run in Container Runtime, a set of preconfigured ML environments hosted on Snowpark Container Services. The compute behind them is a compute pool you size and choose as CPU or GPU, which is separate from a virtual warehouse.
1.Compute pools: the hardware behind Container Runtime
Snowflake gives ML work two kinds of compute: virtual warehouses and compute pools. This page explains both, starting with compute pools because most Snowflake ML infrastructure runs on them. The Snowflake Container Runtime is a set of preconfigured environments for machine learning, built on Snowpark Container Services. Your custom Python ML code and the supported training APIs run inside Snowpark Container Services, and that code needs somewhere to run. That place is a compute pool: a group of nodes that are all the same machine type. ML Jobs, Notebooks on Container Runtime and other Snowpark Container Services workloads all get scheduled onto a pool.
You create a pool with three required properties. MIN_NODES is the minimum number of nodes the pool keeps, and it must be greater than 0. MAX_NODES is the most nodes the pool can scale to. INSTANCE_FAMILY sets the machine type for every node, so it decides whether you get CPU or GPU nodes, how big they are, and how many credits the pool uses while it runs. The optional properties cover how the pool behaves over its lifecycle: AUTO_RESUME, INITIALLY_SUSPENDED and AUTO_SUSPEND_SECS.
CREATE COMPUTE POOL [ IF NOT EXISTS ] <name>
[ FOR APPLICATION <app-name> ]
MIN_NODES = <num>
MAX_NODES = <num>
INSTANCE_FAMILY = <instance_family_name>
[ AUTO_RESUME = { TRUE | FALSE } ]
[ INITIALLY_SUSPENDED = { TRUE | FALSE } ]
[ AUTO_SUSPEND_SECS = <num> ]For ML Jobs, the documentation gives a default pool size: the CPU_X64_S instance family, with at least 1 node and at most 25. If that is too small, or you need GPUs, you create your own pool with a different INSTANCE_FAMILY and pass its name when you submit the job. The ML Jobs documentation says jobs can run on GPU and high-memory CPU instances. Which of those you get depends on the instance family you chose for the pool.
Checkpoint 1 of 4· Check yourself
A team wants to know what decides whether their compute pool has GPUs and how many credits it uses while running. Which property is it?
INSTANCE_FAMILY sets the machine type, which decides the compute resources on each node and therefore the credits used. MAX_NODES only limits how many of those nodes the pool can have.
“determines the amount of compute resources in the compute pool and, therefore, the number of credits consumed”Source: docs.snowflake.com
Checkpoint 2 of 4· Exam question
A team needs a proprietary C++-backed scoring library, which is not in the Snowflake conda channel, available inside a training workload on Snowflake. Which approach supports it with the least departure from Snowflake-native operation?
Correct answer: C — Build a Docker image containing the library, push it to a Snowflake image repository, and run it as a service on a compute pool
- A. Incorrect: the Anaconda channel in Snowflake is a curated set of packages; a proprietary library is not in it, and a notebook cell cannot add it to that channel.
- B. Incorrect: stage imports into Snowpark stored procedures only work for pure-Python code, and a compiled native library cannot be loaded this way on a warehouse.
- C. Correct: a container image bundles arbitrary native dependencies, is pushed to a Snowflake image repository, and runs under Snowpark Container Services on a compute pool, so the workload stays inside Snowflake.
- D. Incorrect: Snowpark-optimized warehouses provide more memory per node, not the ability to install arbitrary packages, and warehouse sandboxes do not permit runtime pip installs of native code.
2.Idle cost and who can operate the pool
Because the instance family sets the credits a pool uses while it is running, a pool that runs with nothing to do still costs money. Two properties, which you can set at creation or change later with ALTER COMPUTE POOL, deal with this. AUTO_SUSPEND_SECS is how many seconds of inactivity Snowflake waits before suspending the pool. Inactivity has a specific meaning here: no services and no jobs running on any node in the pool. AUTO_RESUME sets whether submitting a service or job wakes a suspended pool automatically. If AUTO_RESUME is FALSE, someone has to run ALTER COMPUTE POOL <name> RESUME before any job can start. So for occasional training runs, the main setting to reduce idle cost is a shorter AUTO_SUSPEND_SECS, with AUTO_RESUME left on so the next job still starts without manual help.
Access to a pool is split into separate privileges, and exam scenarios often describe a role that needs two of them. Running a job and suspending a pool are separate rights.
| Privilege | What it allows |
|---|---|
| USAGE | Running a service or a job on the pool |
| OPERATE | Suspending or resuming the pool |
| MODIFY | Altering the pool and setting its properties |
| MONITOR | Viewing usage (services and jobs running), properties, and listing pools |
| OWNERSHIP | Full control; only one role can hold it on a given pool at a time |
Checkpoint 3 of 4· Match them up
Match each task to the compute pool privilege it needs
Tap a term, then the definition that fits it.
USAGE lets a role run jobs, OPERATE lets it suspend and resume, MODIFY lets it change properties, and MONITOR lets it view usage. A data scientist who both runs jobs and suspends the pool needs USAGE and OPERATE.
“OPERATE | Enables suspending or resuming a compute pool.”Source: docs.snowflake.com
Sources3
3.Allocating work: Snowpark-optimized warehouse or Container Runtime
Not every training job needs a compute pool. Snowpark-optimized warehouses are virtual warehouses built for workloads that need a lot of memory and compute. With them, you can train a model with custom code inside a Snowpark Python stored procedure. The procedure runs nested queries to load and transform data, pulls the data into its own memory, trains the model, and can write the trained model to a stage. The documentation describes this pattern as single-node training.
Container Runtime is the option when that single node is not enough. It runs on CPU or GPU compute pools and comes with popular ML and deep learning frameworks already installed. When you use the Snowflake ML APIs, it spreads the processing across the available resources. Its distributed framework uses every GPU on a multi-GPU node by default. You can also set the number of resources for each task, including GPU and memory, through the provided APIs. Container Runtimes are versioned, so you can pin a workload to a specific version and upgrade when you choose.
| Aspect | Snowpark-optimized warehouse | Container Runtime on a compute pool |
|---|---|---|
| How the code runs | Snowpark Python stored procedure | Notebooks and ML Jobs in Snowpark Container Services |
| Scale described in the docs | Single-node ML training | Processing distributed across available resources, including multi-GPU nodes |
| Hardware choice | Memory-heavy warehouse | INSTANCE_FAMILY: CPU or GPU nodes |
| Environment | Stored procedure packages | Versioned, preconfigured ML images; custom images possible |
Checkpoint 4 of 4· Check yourself
A team trains a scikit-learn model in a Python stored procedure. The data fits in memory on one machine, but standard warehouses run out of memory. What is the documented way to give this workload more resources?
Snowpark-optimized warehouses are meant for memory-heavy, single-node training through stored procedures. Compute pools and custom runtimes are Container Runtime tools, and this workload does not need them.
“Snowpark-optimized warehouses make it possible to use Snowpark stored procedures to run single-node ML training workloads directly in Snowflake.”Source: docs.snowflake.com
When the Snowflake-provided Container Runtime image does not contain the libraries or exact versions you need, you can deploy your own runtime on Container Runtime instead of installing packages by hand in every session. You build a custom container image on top of an official Snowflake ML runtime base image, adding your own packages and configuration. Base images are specific to the hardware type (CPU or GPU) and Python version, so pick the one that matches your compute pool. You then push the image to a Snowflake image repository, validate it, and register it as a Custom Runtime Environment (CRE). Custom images also suit regulated teams that need pre-scanned, allowlisted images, and teams without external access to PyPI at runtime, because the packages are already baked into the image.
The workflow has a fixed shape. Log the local Docker client in to the Snowflake image registry with snow spcs image-registry login. Write a Dockerfile that starts FROM the Snowflake base image and installs your packages. Build it for the linux/amd64 platform. Validate it locally with snow custom-image validate, which checks for the exact entrypoint /usr/local/bin/entrypoint.sh, the DASHBOARD_PORT=12003 environment variable, a set of mandatory Python packages, and a conflict-free dependency tree. Vulnerability scanning is optional. Then tag and push the image to your image repository, and run the registration DDL. Snowflake runs a server-side validation during registration, and the CRE is created only if that check passes.
CREATE CUSTOM RUNTIME ENVIRONMENT my_cre
IMAGE_PATH = '/<database>/<schema>/<image_repo>/my-custom-image:v1'
BASE_IMAGE_TYPE = CPU;Once the CRE is in the RESOLVED state, you reference it from an ML Job with the runtime_environment parameter in the form cre@<name>, for example runtime_environment="cre@my_cre" on the @remote decorator. The same keyword is supported by submit_file, submit_directory and submit_from_stage. The active role needs the USAGE privilege on the CRE plus access to the underlying image repository, and a DEPRECATED CRE cannot be used for new jobs. The system also checks the image's SHA digest at runtime: if the image in the repository was modified after the CRE was registered, the job submission is rejected and you must recreate the CRE. Notebooks reference a CRE the same way, using cre@<name>. If you leave runtime_environment out, Snowflake uses the latest Snowflake Container Runtime version on your compute pool.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.USAGE on a compute pool is enough to suspend and resume it manually.Why is that wrong?
USAGE only lets a role run services and jobs on the pool. Suspending and resuming need the separate OPERATE privilege.
Covered in Idle cost and who can operate the pool
2.A custom runtime image can be built from any base image, such as a generic Python image, as long as it contains your libraries.Why is that wrong?
Custom images must derive from an official Snowflake ML runtime base image, matched to the hardware type and Python version, and must pass validation before registration.
Covered in Allocating work: Snowpark-optimized warehouse or Container Runtime
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The default compute pool size uses the CPU_X64_S instance family. The minimum number of nodes is 1 and the maximum is 25.”
↩︎ Compute pools: the hardware behind Container Runtime“Run ML workloads on Snowflake Compute Pools, including GPU and high-memory CPU instances.”
↩︎ Compute pools: the hardware behind Container Runtime - 2.
“Specifies the minimum number of nodes for the compute pool. This value must be greater than 0.”
↩︎ Compute pools: the hardware behind Container Runtime“determines the amount of compute resources in the compute pool and, therefore, the number of credits consumed”
↩︎ Checkpoint - 3.
“Inactivity means no services and no jobs running on any node in the compute pool.”
↩︎ Idle cost and who can operate the pool“Specifies whether to automatically resume a compute pool when a service or job is submitted to it.”
↩︎ Idle cost and who can operate the pool“Number of seconds of inactivity after which you want Snowflake to automatically suspend the compute pool.”
↩︎ Prediction - 4.https://docs.snowflake.com/en/developer-guide/snowpark/python/python-snowpark-training-mlOfficial docs
“you can use them to train an ML model using custom code on a single node.”
↩︎ Allocating work: Snowpark-optimized warehouse or Container Runtime“Snowpark-optimized warehouses make it possible to use Snowpark stored procedures to run single-node ML training workloads directly in Snowflake.”
↩︎ Checkpoint - 5.
“By default, this framework uses all GPUs on multi-GPU nodes”
↩︎ Allocating work: Snowpark-optimized warehouse or Container Runtime“When using the Snowflake ML APIs, the Container Runtime distributes the processing across available resources.”
↩︎ Allocating work: Snowpark-optimized warehouse or Container Runtime“which offers the ability to run on CPU or GPU compute pools”
↩︎ Key concept - 6.
“After building a custom image, you push it to a Snowflake image repository, validate it, and register it as a Custom Runtime Environment (CRE).”
↩︎ Allocating work: Snowpark-optimized warehouse or Container Runtime“All custom images must be derived from an official Snowflake ML runtime base image.”
↩︎ Allocating work: Snowpark-optimized warehouse or Container Runtime“Reference an admin-approved Custom Runtime Environment using the cre@<name> format.”
↩︎ Allocating work: Snowpark-optimized warehouse or Container Runtime“All custom images must be derived from an official Snowflake ML runtime base image.”
↩︎ Exam trap 2
Also cited
“USAGE | Enables running a service or a job.”
↩︎ Exam trap 1“OPERATE | Enables suspending or resuming a compute pool.”
↩︎ Checkpoint