What you will be able to do
- Choose between resizing, a Snowpark-optimized warehouse and other warehouse strategies for an ML workload
- Explain per-second warehouse billing and its 60-second minimum
- Size a compute pool with an instance family and MIN_NODES/MAX_NODES, and use AUTO_SUSPEND
- Track compute pool credits and other SPCS costs in the ACCOUNT_USAGE and ORGANIZATION_USAGE views
1.Warehouse sizing and billing for ML workloads
Warehouse credits depend on three things: how many warehouses run, for how long, and at what size. Moving up one size roughly doubles both the compute and the hourly credit rate. Upsizing therefore only pays off if the job finishes in about half the time or less. Billing is per second with a 60-second minimum each time a warehouse starts or resumes. Suspended warehouses consume no credits. A warehouse that suspends and resumes within the first minute is charged again, because the 60-second minimum restarts on each resume.
The performance guidance lists several warehouse strategies: reduce queues, resolve memory spillage, increase warehouse size, try query acceleration, optimize the warehouse cache, and limit concurrent queries. Memory spillage is the one that matters most for ML. When a warehouse runs out of memory, bytes spill onto storage and the query slows down considerably. A training stored procedure that runs on a single node is limited by memory rather than parallelism, and adding nodes does not help it.
For that case, Snowflake offers Snowpark-optimized warehouses. Their default configuration provides 16x the memory per node of a standard warehouse. RESOURCE_CONSTRAINT selects the memory and CPU architecture, and higher memory tiers require a minimum warehouse size. These warehouses can take longer to create and resume, and workloads that don't use Snowpark may not benefit from them.
| Memory (up to) | RESOURCE_CONSTRAINT values | Minimum warehouse size |
|---|---|---|
| 16GB | MEMORY_1X, MEMORY_1X_x86 | XSMALL |
| 256GB | MEMORY_16X, MEMORY_16X_x86 | M |
| 1TB | MEMORY_64X, MEMORY_64X_x86 | L |
CREATE OR REPLACE WAREHOUSE snowpark_opt_wh WITH
WAREHOUSE_SIZE = 'MEDIUM'
WAREHOUSE_TYPE = 'SNOWPARK-OPTIMIZED';Tuning is easier when each warehouse runs similar work. If a warehouse runs very different queries, the cost of a performance enhancement may be spent on queries that don't benefit from it. This is another reason to give ML training its own warehouse, and that separate warehouse can also carry its own cost tag.
Checkpoint 1 of 5· Check yourself
A single-node Snowpark training stored procedure on a standard warehouse spills heavily to storage. Which change is the documented fit?
Snowpark-optimized warehouses add memory per node, which addresses spillage in memory-bound Snowpark work such as single-node ML training.
“Snowpark-optimized warehouses are recommended for running Snowpark workloads such as code that has large memory requirements or dependencies on a specific CPU architecture.”Source: docs.snowflake.com
Checkpoint 2 of 5· Exam question
A team sums `CREDITS_ATTRIBUTED_COMPUTE` from QUERY_ATTRIBUTION_HISTORY for a month on its ML warehouse and finds the total noticeably lower than the warehouse's credits in WAREHOUSE_METERING_HISTORY. What explains the gap?
Correct answer: B — Per-query attribution excludes warehouse idle time, so credits burned while the warehouse ran with no active queries appear only in the metering history.
- A. Incorrect: short queries are still attributed their share of compute; rounding to zero is not how the view works.
- B. Correct: attributed compute covers only time queries actually ran, so idle time (such as the auto-suspend wait) is missing and explains the gap.
- C. Incorrect: the view covers queries from all roles in the account, not just the administrator role, and all warehouse usage is billed as warehouse credits.
- D. Incorrect: both views report credits in the same unit; no size-factor conversion is applied to reconcile them.
2.Sizing a compute pool: instance family and node limits
A compute pool is the SPCS counterpart of a warehouse. It is an account-level collection of VM nodes. Instead of a size, you choose an instance family, which sets each node's vCPU, memory, storage and egress bandwidth, including whether the nodes have GPUs. You also set a minimum and maximum node count, and Snowflake autoscales between them. Cost follows directly from these settings: the number and type of nodes determine the credits consumed. Run SHOW COMPUTE POOL INSTANCE FAMILIES to see which families are available in your region.
CREATE COMPUTE POOL tutorial_compute_pool MIN_NODES = 1 MAX_NODES = 1 INSTANCE_FAMILY = CPU_X64_XS;The two node limits work in opposite directions. Raising MIN_NODES above 1 keeps nodes warm for bursty inference traffic, so you don't wait for autoscaling, but you pay for those nodes even when they are idle. MAX_NODES is a cost guardrail: it caps how far autoscaling can grow during a load spike or when a code bug requests more nodes than planned. Right-sizing the instance family has the biggest effect. A GPU family on a small model that barely uses the GPU costs GPU credits for CPU-sized work.
| Setting | What it controls | Cost effect |
|---|---|---|
| INSTANCE_FAMILY | Machine type of each node | Sets the credit rate per node |
| MIN_NODES | Nodes the pool launches with | Higher values keep capacity ready but billed |
| MAX_NODES | Ceiling for Snowflake autoscaling | Caps runaway node growth |
| AUTO_RESUME | Start a suspended pool when a service or job arrives | Allows the pool to stay suspended when unused |
Checkpoint 3 of 5· Match them up
Match each compute pool setting to its role
Tap a term, then the definition that fits it.
The instance family sets the per-node rate, MAX_NODES caps growth, and a MIN_NODES value above 1 trades idle cost for readiness.
“Setting a maximum node limit prevents an unexpectedly large number of nodes from being added to your compute pool by Snowflake autoscaling.”Source: docs.snowflake.com
3.Tracking and trimming SPCS compute pool costs
A compute pool is billed whenever it is IDLE, ACTIVE, STOPPING or RESIZING. It is not billed while STARTING or SUSPENDED. A pool that sits idle overnight therefore keeps costing money, which is why the documentation's main optimization advice is to use AUTO_SUSPEND. AUTO_RESUME then starts the pool again when a service is created or called.
To track pool credits, use ACCOUNT_USAGE.SNOWPARK_CONTAINER_SERVICES_HISTORY. It gives hourly credits per pool for the last 365 days, with up to 3 hours of latency. Because each row carries COMPUTE_POOL_NAME and CREDITS_USED, giving each team its own pool yields per-team credits directly. IS_EXCLUSIVE and APPLICATION_NAME identify pools created for an application. For combined reports, filter METERING_HISTORY or METERING_DAILY_HISTORY on service_type SNOWPARK_CONTAINER_SERVICES. At organization level, use the same filter on ORGANIZATION_USAGE METERING_DAILY_HISTORY.
Compute is only one of three SPCS cost categories; the other two are storage and data transfer. Storage costs include the image repository (a stage), event-table logs, mounted stage volumes and block storage. Local node storage mounted as a volume costs nothing extra. Data transfer covers egress to other regions or the internet, which appears in DATA_TRANSFER_HISTORY under transfer_type SNOWPARK_CONTAINER_SERVICES. It also covers internal transfer between compute pools and warehouses caused by service functions, which appears in INTERNAL_DATA_TRANSFER_HISTORY. Data transfer is currently not billed on Google Cloud accounts.
Checkpoint 4 of 5· Check yourself
Finance wants monthly credits per ML team, and each team trains on its own compute pool. Which view gives this most directly?
This view reports hourly credits for each compute pool. With one pool per team, summing by pool name gives per-team credits.
“The SNOWPARK_CONTAINER_SERVICES_HISTORY view offers credit usage information (hourly consumption) exclusively for Snowpark Container Services.”Source: docs.snowflake.com
Checkpoint 5 of 5· Exam question
A single application submits inference queries on behalf of marketing, risk, and support users, all using one application service user and one warehouse. Which technique lets cost reports separate the departments?
Correct answer: A — Set the `QUERY_TAG` session parameter to the requester's cost center before each query, then group QUERY_ATTRIBUTION_HISTORY credits by `QUERY_TAG`.
- A. Correct: a per-request QUERY_TAG travels with each query and QUERY_ATTRIBUTION_HISTORY exposes it, so one user and one warehouse can still be split by cost center.
- B. Incorrect: WAREHOUSE_METERING_HISTORY is hourly per warehouse and has no role column, so it cannot be grouped by role.
- C. Incorrect: a resource monitor is assigned to a warehouse or account and tracks that scope's credits; it does not split usage by department.
- D. Incorrect: a tag on the single service user labels all queries with one value, so it cannot distinguish departments.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Suspending a warehouse quickly between short ML steps always saves credits.Why is that wrong?
Each resume is billed for at least one minute, so suspending and resuming within the first minute leads to repeated minimum charges.
2.A compute pool with no active services costs nothing.Why is that wrong?
IDLE is a billed state. Only STARTING and SUSPENDED are free, so idle pools should use AUTO_SUSPEND.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“approximately doubles the computing power and the number of credits billed per full hour that the warehouse runs”
↩︎ Warehouse sizing and billing for ML workloads“credits are billed per-second, with a 60-second (i.e. 1-minute) minimum”
↩︎ Warehouse sizing and billing for ML workloads“results in multiple charges because the 1-minute minimum starts over each time a warehouse is resumed”
↩︎ Exam trap 1 - 2.
“a query runs substantially slower when a warehouse runs out of memory”
↩︎ Warehouse sizing and billing for ML workloads“the cost of a performance enhancement might be wasted on a query that does not benefit from the optimization”
↩︎ Warehouse sizing and billing for ML workloads - 3.
“The default configuration for a Snowpark-optimized warehouse provides 16x memory per node compared to a standard warehouse.”
↩︎ Warehouse sizing and billing for ML workloads“Snowpark-optimized warehouses are recommended for running Snowpark workloads such as code that has large memory requirements or dependencies on a specific CPU architecture.”
↩︎ Checkpoint - 4.https://docs.snowflake.com/en/developer-guide/snowpark-container-services/working-with-compute-poolOfficial docs
“Specifying an instance family when creating a compute pool is similar to specifying warehouse size”
↩︎ Sizing a compute pool: instance family and node limits“This approach ensures that additional nodes are readily available when needed, instead of waiting for autoscaling to start.”
↩︎ Sizing a compute pool: instance family and node limits“Setting a maximum node limit prevents an unexpectedly large number of nodes from being added to your compute pool by Snowflake autoscaling.”
↩︎ Checkpoint - 5.https://docs.snowflake.com/en/developer-guide/snowpark-container-services/accounts-orgs-usage-viewsOfficial docs
“The number and type (instance family) of the nodes in the compute pool (see CREATE COMPUTE POOL) determine the credits it consumes”
↩︎ Sizing a compute pool: instance family and node limits“but not when it is in a STARTING or SUSPENDED state”
↩︎ Tracking and trimming SPCS compute pool costs“The costs associated with using Snowpark Container Services can be categorized into storage cost, compute pool cost, and data transfer cost.”
↩︎ Tracking and trimming SPCS compute pool costs“Data transfer costs are currently not billed for Snowflake accounts on Google Cloud.”
↩︎ Tracking and trimming SPCS compute pool costs“To optimize compute pool expenses, you should leverage the AUTO_SUSPEND feature”
↩︎ Exam trap 2“You incur charges for a compute pool in the IDLE, ACTIVE, STOPPING, or RESIZING state”
↩︎ Prediction“The SNOWPARK_CONTAINER_SERVICES_HISTORY view offers credit usage information (hourly consumption) exclusively for Snowpark Container Services.”
↩︎ Checkpoint - 6.https://docs.snowflake.com/en/sql-reference/account-usage/snowpark_container_services_historyOfficial docs
“Latency for the view may be up to 180 minutes (3 hours).”
↩︎ Tracking and trimming SPCS compute pool costs