What you will be able to do
- Submit ML Jobs with @remote, submit_file, submit_directory or submit_from_stage, then monitor them and get their results
- Run a multi-node ML Job with target_instances and min_instances on a compute pool with enough nodes
- Choose between open-source training and the snowflake.ml.modeling.distributors estimators
- Build, validate, register and reference a Custom Runtime Environment for packages not in the base image
1.Operating ML Jobs
Snowflake ML Jobs run ML workflows in Container Runtime on a compute pool, and you can start them from any development environment, such as VS Code or Jupyter. They need snowflake-ml-python 1.26.0 or later and an active Snowpark Session. Without a session, even list_jobs() fails. There are two ways to submit a job. Function dispatch uses the @remote decorator: it serializes the function, uploads it to a stage, and runs it in Container Runtime. File dispatch has three entry points. submit_file runs a single Python file. submit_directory runs a project spread across several files. submit_from_stage runs a project that is already stored on a stage. File dispatch also accepts command-line arguments, environment variables and extra imports.
Every submission method accepts a runtime_environment keyword. If you leave it out, the job runs on the newest Container Runtime available on the compute pool, so its environment can change between runs. If you set it to a version string such as 2.3.0, the job is pinned to that version.
Checkpoint 1 of 6· Fill the gap
Which keyword pins this function to Container Runtime version 2.3.0?
from snowflake.ml.jobs import remote
@remote("MY_COMPUTE_POOL", stage_name="payload_stage", session=session, ? ="2.3.0")@remote, submit_file, submit_directory and submit_from_stage all take runtime_environment. If it is omitted, the latest available runtime on the pool is used.
Source: docs.snowflake.comEvery submission returns an MLJob object, which you use to manage and monitor the job. A decorated function returns its result through result(). A file-based job returns a value by assigning it to the special __return__ variable. result() waits until the job finishes, then returns the value, or raises an exception if the job failed. Inside the job, you can get a Snowpark Session with Session.builder.getOrCreate().
from snowflake.ml.jobs import MLJob, get_job, list_jobs
# List all jobs
jobs = list_jobs()
# Retrieve an existing job based on ID
job = get_job("<job_id>") # job is an MLJob instance
# Basic job information
print(f"Job ID: {job.id}")
print(f"Status: {job.status}") # PENDING, RUNNING, FAILED, DONE
# Wait for completion
job.wait()To read a job's output, call get_logs() on the MLJob, or show_logs() to print the logs. For a multi-node job both default to the head instance, and you can pass instance_id to read the logs of a specific instance. The overview points to the Ray Dashboard in ML Jobs for further management and monitoring of a job.
Checkpoint 2 of 6· Exam question
Which TWO steps are required to make a custom container image usable by a service in Snowpark Container Services?(Select 2)
Correct answers: D, E — Reference the pushed image by its repository path in the service specification's container image field, so the pool pulls it; Build the image outside Snowflake, create a repository with CREATE IMAGE REPOSITORY, and push the image to it with Docker
- A. Incorrect: Snowflake does not build images from a Dockerfile in a stage; CREATE SERVICE only references an already built image in the registry.
- B. Incorrect: warehouses do not run user containers and have no image setting; containers run only on compute pools.
- C. Incorrect: external volumes point at cloud object storage for data, not container registries, and they play no role in image distribution.
- D. Correct: the service specification's image field must point at the pushed image path in the repository so that the compute pool pulls the right image.
- E. Correct: Snowflake offers an OCI-compliant registry; images are built outside Snowflake and pushed into an image repository created with CREATE IMAGE REPOSITORY.
2.Distributed training: multi-node jobs and the distributors API
A single-node ML Job has only a head node. To spread a job across several nodes, pass target_instances; multi-node jobs need snowflake-ml-python 1.9.2 or later. One node coordinates as the head node and the others work as worker nodes, but together they appear as one job. The compute pool has to allow that many nodes: MAX_NODES must be greater than or equal to target_instances. Before running your code, ML Jobs wait until target_instances nodes are available. If they do not all arrive before the timeout, the job fails. If you also set min_instances, the job starts as soon as that smaller number of nodes is ready.
from snowflake.ml.jobs import submit_file
job = submit_file(
"<script_path>",
"MY_COMPUTE_POOL",
stage_name="<payload_stage>",
session=session,
target_instances=<num_training_nodes> # Specify the number of nodes
)Extra nodes only help if the code is written to use them, through Ray or Snowflake's Distributed Modeling Classes. Container Runtime offers distributed versions of XGBoost (XGBEstimator), LightGBM (LightGBMEstimator) and PyTorch (PyTorchDistributor) in the snowflake.ml.modeling.distributors namespace. You can pass a DataConnector straight to their fit method. A scaling config such as XGBScalingConfig(use_gpu=True) is optional, because by default they use all available resources. Comparing the two environments: the sources present these distributed APIs as part of Container Runtime. For warehouses, they describe Snowpark-optimized warehouses running single-node training in stored procedures, and they do not describe a distributed equivalent there, so treat that as what the material covers rather than a statement about every possible configuration.
| Approach | Runs on | Use it when |
|---|---|---|
| Stored procedure on a Snowpark-optimized warehouse | Virtual warehouse | Custom training code fits on a single node |
| Open-source XGBoost, LightGBM or PyTorch | Container Runtime | Small datasets, rapid prototyping, lift-and-shift without distributed needs |
| XGBEstimator, LightGBMEstimator, PyTorchDistributor | Container Runtime, multi-node or multi-GPU | Data larger than one node's memory, or several GPUs to use |
Checkpoint 3 of 6· Check yourself
A job sets target_instances=5 and min_instances=3. Only 3 nodes are available at first. What happens?
min_instances lets the job start with fewer nodes than target_instances. Without it, the job waits for all target nodes and fails if they do not arrive before the timeout.
“the job payload is executed as soon as the minimum number of nodes becomes available”Source: docs.snowflake.com
Checkpoint 4 of 6· Exam question
A nightly model training run takes four hours and must not be restarted automatically if the process exits with an error, because partial runs corrupt an output table. Which TWO choices fit?(Select 2)
Correct answers: C, D — Run it on a compute pool that has a short AUTO_SUSPEND_SECS so that the pool suspends once the job finishes; A job service started with EXECUTE JOB SERVICE, because exited containers are not restarted and the job has a finite lifetime
- A. Incorrect: regular services are intended for long-running applications and Snowflake restarts containers that exit, which would re-run the partial training.
- B. Incorrect: a service is not dropped automatically when its container exits; it stays and keeps being restarted until explicitly dropped or suspended.
- C. Correct: a short auto-suspend lets the pool release nodes after the finite job ends, so the nightly run is not billed for idle time.
- D. Correct: job services run to completion like stored procedures, and Snowflake does not restart exited job containers, which prevents repeated partial writes.
- E. Incorrect: service functions expose a service endpoint to SQL for inference-style calls and are not a batch training construct with run-once semantics.
3.Custom runtime images for libraries the base image lacks
Installing packages with pip at runtime is not always possible or allowed. For example, a regulated team may require pre-scanned images, or an account may have no external access integration to PyPI. Custom runtime images solve this. Every custom image has to be built on an official Snowflake ML runtime base image, and those base images are specific to hardware (CPU or GPU) and Python version. You add your packages on top of the base image, build for linux/amd64, and validate the image locally with snow custom-image validate.
FROM <registry_url>/snowflake/images/snowflake_images/container_runtime/cpu_x86_64:2.7.2
RUN uv pip install --system --break-system-packages \
xgboost==2.0.3 \
scikit-learn==1.3.2To pass validation, the image must use the entrypoint /usr/local/bin/entrypoint.sh, set DASHBOARD_PORT=12003, include the required packages (snowflake-ml-python, snowflake-snowpark-python[pandas], ray, jupyter-server and others), and have no dependency conflicts. Vulnerability scanning with --scan-vulnerabilities is optional. Once it validates, tag the image and push it to a Snowflake image repository. Then register it with CREATE CUSTOM RUNTIME ENVIRONMENT. Snowflake runs its own validation at this step and creates the CRE only if that check passes, leaving it in the RESOLVED state.
CREATE CUSTOM RUNTIME ENVIRONMENT my_cre
IMAGE_PATH = '/<database>/<schema>/<image_repo>/my-custom-image:v1'
BASE_IMAGE_TYPE = CPU;To use the CRE, ML Jobs set runtime_environment="cre@my_cre", which requires snowflake-ml-python 1.37.0 or later. Notebooks can select it in advanced settings or pass RUNTIME = 'cre@...' when executed from SQL. The role needs USAGE on the CRE and access to the image repository. A CRE must be in the RESOLVED state; a DEPRECATED CRE cannot be used for new work. At submission time, Snowflake checks the image's SHA digest. If the image was changed after the CRE was registered, the submission is rejected and you have to recreate the CRE.
Checkpoint 5 of 6· Check yourself
A Custom Runtime Environment is in the DEPRECATED state. What can you do with it?
Only RESOLVED environments can be referenced. DEPRECATED ones can't be used for new jobs (or new Notebooks).
“Environments marked as DEPRECATED can’t be used for new jobs.”Source: docs.snowflake.com
Checkpoint 6 of 6· Put it in order
Put the custom runtime image workflow in order
- 1.Pull the matching Snowflake base image
- 2.Authenticate Docker with snow spcs image-registry login
- 3.Tag and push the image to the Snowflake image repository
- 4.Write a Dockerfile based on it and build the image for linux/amd64
- 5.Validate locally with snow custom-image validate
- 6.Run CREATE CUSTOM RUNTIME ENVIRONMENT
You authenticate before pulling or pushing. The Dockerfile is written on the base image and built next. Local validation comes before the push, and registration comes last because it validates the image already stored in the repository.
“Before pushing the image to Snowflake, validate it locally to ensure it meets runtime requirements.”Source: docs.snowflake.com
Sources6
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Leaving out runtime_environment keeps an ML Job on the same runtime version every time.Why is that wrong?
Without the keyword, Snowflake picks the latest Container Runtime available on the pool. To pin a version, pass it explicitly.
Covered in Operating ML Jobs
2.A multi-node job can ask for more nodes than the compute pool's MAX_NODES, and the pool will scale past it.Why is that wrong?
MAX_NODES must be at least the number of target instances the job requests.
Covered in Distributed training: multi-node jobs and the distributors API
3.Pushing an updated image under the same tag automatically updates an existing CRE.Why is that wrong?
Snowflake checks the image's SHA digest at runtime and rejects the submission if the image changed after registration. You have to recreate the CRE.
Covered in Custom runtime images for libraries the base image lacks
Practise it for real
Package a pinned XGBoost version into a Custom Runtime Environment and run an ML Job on it
1.Run snow spcs image-registry login, then pull the cpu_x86_64:2.7.2 base image with docker pull.
Why: Every custom image has to be built on an official Snowflake ML base image, and Docker needs to be logged in to the registry first.
You should see: The base image is in your local Docker image list.
2.Write the Dockerfile from this page and run docker build --platform linux/amd64 -t my-custom-image:v1 .
Why: The image has to be built for linux/amd64 to run in Snowflake's execution environment.
You should see: The local image my-custom-image:v1 exists.
3.Run snow custom-image validate my-custom-image:v1 and fix anything it reports.
Why: Checking locally catches entrypoint, environment variable and dependency problems before the server-side registration check.
You should see: The validation report passes.
4.Tag and push the image to <registry_url>/<database>/<schema>/<image_repo>/my-custom-image:v1, then run CREATE CUSTOM RUNTIME ENVIRONMENT my_cre.
Why: Registering the image runs server-side validation and creates an object you can grant USAGE on.
You should see: my_cre is in the RESOLVED state.
5.Decorate a function with @jobs.remote("MY_COMPUTE_POOL", stage_name="payload_stage", runtime_environment="cre@my_cre") and call it.
Why: The cre@<name> format points the ML Job at your registered environment.
You should see: An MLJob object is returned and its status eventually reaches DONE.
Stuck? Get a nudge
If the job is rejected, check whether the image tag was pushed again after you created the CRE.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Invoking a @remote decorated function returns a Snowflake MLJob object that can be used to manage and monitor the job execution.”
↩︎ Operating ML Jobs“submit_from_stage: For running Python projects saved on a Snowflake stage”
↩︎ Operating ML Jobs“For file-based jobs, use the special __return__ variable to specify the return value.”
↩︎ Operating ML Jobs“Snowflake automatically uses the latest available version of the Snowflake Container Runtime on your compute pool.”
↩︎ Exam trap 1 - 2.
“In multi-node jobs, you can access logs from specific instances:”
↩︎ Operating ML Jobs“ML Jobs automatically wait for the specified target_instances to be available before executing your payload.”
↩︎ Distributed training: multi-node jobs and the distributors API“snowflake-ml-python>=1.9.2”
↩︎ Distributed training: multi-node jobs and the distributors API“You must set MAX_NODES to be greater than or equal to the number of target instances”
↩︎ Exam trap 2“the job payload is executed as soon as the minimum number of nodes becomes available”
↩︎ Checkpoint - 3.
“These are found in the snowflake.ml.modeling.distributors namespace.”
↩︎ Distributed training: multi-node jobs and the distributors API - 4.
“Train models on datasets that are larger than the memory of a single compute node”
↩︎ Distributed training: multi-node jobs and the distributors API“However, Snowflake defaults to using all available resources.”
↩︎ Distributed training: multi-node jobs and the distributors API - 5.https://docs.snowflake.com/en/developer-guide/snowpark/python/python-snowpark-training-mlOfficial docs
“Snowpark-optimized warehouses make it possible to use Snowpark stored procedures to run single-node ML training workloads directly in Snowflake.”
↩︎ Distributed training: multi-node jobs and the distributors API - 6.
“All custom images must be derived from an official Snowflake ML runtime base image.”
↩︎ Custom runtime images for libraries the base image lacks“During this step, Snowflake performs a server-side validation check to confirm the image executes correctly.”
↩︎ Custom runtime images for libraries the base image lacks“The active role must hold the USAGE privilege on the target Custom Runtime Environment, along with access to the underlying image repository.”
↩︎ Custom runtime images for libraries the base image lacks“If the underlying image in the repository was modified after the CRE was registered, the job submission is rejected to ensure reproducibility.”
↩︎ Exam trap 3“Environments marked as DEPRECATED can’t be used for new jobs.”
↩︎ Checkpoint“Before pushing the image to Snowflake, validate it locally to ensure it meets runtime requirements.”
↩︎ Checkpoint