CertSafari
    Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01)· Lessons

    Domain 2 · Lesson 6/17

    ML Jobs, Distributed Training and Custom Runtime Images

    Manage infrastructure for ML.

    11 min read
    8% of exam
    6 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Submit ML Jobs with @remote, submit_file, submit_directory or submit_from_stage, then monitor them and get their results
    • Run a multi-node ML Job with target_instances and min_instances on a compute pool with enough nodes
    • Choose between open-source training and the snowflake.ml.modeling.distributors estimators
    • Build, validate, register and reference a Custom Runtime Environment for packages not in the base image

    1.Operating ML Jobs

    Snowflake ML Jobs run ML workflows in Container Runtime on a compute pool, and you can start them from any development environment, such as VS Code or Jupyter. They need snowflake-ml-python 1.26.0 or later and an active Snowpark Session. Without a session, even list_jobs() fails. There are two ways to submit a job. Function dispatch uses the @remote decorator: it serializes the function, uploads it to a stage, and runs it in Container Runtime. File dispatch has three entry points. submit_file runs a single Python file. submit_directory runs a project spread across several files. submit_from_stage runs a project that is already stored on a stage. File dispatch also accepts command-line arguments, environment variables and extra imports.

    Every submission method accepts a runtime_environment keyword. If you leave it out, the job runs on the newest Container Runtime available on the compute pool, so its environment can change between runs. If you set it to a version string such as 2.3.0, the job is pinned to that version.

    Checkpoint 1 of 6· Fill the gap

    Which keyword pins this function to Container Runtime version 2.3.0?

    from snowflake.ml.jobs import remote
    
    @remote("MY_COMPUTE_POOL", stage_name="payload_stage", session=session,  ? ="2.3.0")

    Every submission returns an MLJob object, which you use to manage and monitor the job. A decorated function returns its result through result(). A file-based job returns a value by assigning it to the special __return__ variable. result() waits until the job finishes, then returns the value, or raises an exception if the job failed. Inside the job, you can get a Snowpark Session with Session.builder.getOrCreate().

    Listing, retrieving and waiting on ML Jobspython
    from snowflake.ml.jobs import MLJob, get_job, list_jobs
    
    # List all jobs
    jobs = list_jobs()
    
    # Retrieve an existing job based on ID
    job = get_job("<job_id>")  # job is an MLJob instance
    
    # Basic job information
    print(f"Job ID: {job.id}")
    print(f"Status: {job.status}")  # PENDING, RUNNING, FAILED, DONE
    
    # Wait for completion
    job.wait()

    To read a job's output, call get_logs() on the MLJob, or show_logs() to print the logs. For a multi-node job both default to the head instance, and you can pass instance_id to read the logs of a specific instance. The overview points to the Ray Dashboard in ML Jobs for further management and monitoring of a job.

    Checkpoint 2 of 6· Exam question

    Which TWO steps are required to make a custom container image usable by a service in Snowpark Container Services?(Select 2)

    Sources12

    2.Distributed training: multi-node jobs and the distributors API

    A single-node ML Job has only a head node. To spread a job across several nodes, pass target_instances; multi-node jobs need snowflake-ml-python 1.9.2 or later. One node coordinates as the head node and the others work as worker nodes, but together they appear as one job. The compute pool has to allow that many nodes: MAX_NODES must be greater than or equal to target_instances. Before running your code, ML Jobs wait until target_instances nodes are available. If they do not all arrive before the timeout, the job fails. If you also set min_instances, the job starts as soon as that smaller number of nodes is ready.

    Submitting a file as a multi-node ML Jobpython
    from snowflake.ml.jobs import submit_file
    
    job = submit_file(
        "<script_path>",
        "MY_COMPUTE_POOL",
        stage_name="<payload_stage>",
        session=session,
        target_instances=<num_training_nodes>  # Specify the number of nodes
    )

    Extra nodes only help if the code is written to use them, through Ray or Snowflake's Distributed Modeling Classes. Container Runtime offers distributed versions of XGBoost (XGBEstimator), LightGBM (LightGBMEstimator) and PyTorch (PyTorchDistributor) in the snowflake.ml.modeling.distributors namespace. You can pass a DataConnector straight to their fit method. A scaling config such as XGBScalingConfig(use_gpu=True) is optional, because by default they use all available resources. Comparing the two environments: the sources present these distributed APIs as part of Container Runtime. For warehouses, they describe Snowpark-optimized warehouses running single-node training in stored procedures, and they do not describe a distributed equivalent there, so treat that as what the material covers rather than a statement about every possible configuration.

    Choosing a training approach
    ApproachRuns onUse it when
    Stored procedure on a Snowpark-optimized warehouseVirtual warehouseCustom training code fits on a single node
    Open-source XGBoost, LightGBM or PyTorchContainer RuntimeSmall datasets, rapid prototyping, lift-and-shift without distributed needs
    XGBEstimator, LightGBMEstimator, PyTorchDistributorContainer Runtime, multi-node or multi-GPUData larger than one node's memory, or several GPUs to use

    Checkpoint 3 of 6· Check yourself

    A job sets target_instances=5 and min_instances=3. Only 3 nodes are available at first. What happens?

    Checkpoint 4 of 6· Exam question

    A nightly model training run takes four hours and must not be restarted automatically if the process exits with an error, because partial runs corrupt an output table. Which TWO choices fit?(Select 2)

    Sources2345

    3.Custom runtime images for libraries the base image lacks

    Installing packages with pip at runtime is not always possible or allowed. For example, a regulated team may require pre-scanned images, or an account may have no external access integration to PyPI. Custom runtime images solve this. Every custom image has to be built on an official Snowflake ML runtime base image, and those base images are specific to hardware (CPU or GPU) and Python version. You add your packages on top of the base image, build for linux/amd64, and validate the image locally with snow custom-image validate.

    Dockerfile that extends the CPU base image with pinned librariesdockerfile
    FROM <registry_url>/snowflake/images/snowflake_images/container_runtime/cpu_x86_64:2.7.2
    
    RUN uv pip install --system --break-system-packages \
        xgboost==2.0.3 \
        scikit-learn==1.3.2

    To pass validation, the image must use the entrypoint /usr/local/bin/entrypoint.sh, set DASHBOARD_PORT=12003, include the required packages (snowflake-ml-python, snowflake-snowpark-python[pandas], ray, jupyter-server and others), and have no dependency conflicts. Vulnerability scanning with --scan-vulnerabilities is optional. Once it validates, tag the image and push it to a Snowflake image repository. Then register it with CREATE CUSTOM RUNTIME ENVIRONMENT. Snowflake runs its own validation at this step and creates the CRE only if that check passes, leaving it in the RESOLVED state.

    Registering a Custom Runtime Environmentsql
    CREATE CUSTOM RUNTIME ENVIRONMENT my_cre
        IMAGE_PATH = '/<database>/<schema>/<image_repo>/my-custom-image:v1'
        BASE_IMAGE_TYPE = CPU;

    To use the CRE, ML Jobs set runtime_environment="cre@my_cre", which requires snowflake-ml-python 1.37.0 or later. Notebooks can select it in advanced settings or pass RUNTIME = 'cre@...' when executed from SQL. The role needs USAGE on the CRE and access to the image repository. A CRE must be in the RESOLVED state; a DEPRECATED CRE cannot be used for new work. At submission time, Snowflake checks the image's SHA digest. If the image was changed after the CRE was registered, the submission is rejected and you have to recreate the CRE.

    Checkpoint 5 of 6· Check yourself

    A Custom Runtime Environment is in the DEPRECATED state. What can you do with it?

    Checkpoint 6 of 6· Put it in order

    Put the custom runtime image workflow in order

    1. 1.Pull the matching Snowflake base image
    2. 2.Authenticate Docker with snow spcs image-registry login
    3. 3.Tag and push the image to the Snowflake image repository
    4. 4.Write a Dockerfile based on it and build the image for linux/amd64
    5. 5.Validate locally with snow custom-image validate
    6. 6.Run CREATE CUSTOM RUNTIME ENVIRONMENT

    Sources6

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Leaving out runtime_environment keeps an ML Job on the same runtime version every time.Why is that wrong?

      Without the keyword, Snowflake picks the latest Container Runtime available on the pool. To pin a version, pass it explicitly.

      Covered in Operating ML Jobs

    2. 2.A multi-node job can ask for more nodes than the compute pool's MAX_NODES, and the pool will scale past it.Why is that wrong?

      MAX_NODES must be at least the number of target instances the job requests.

      Covered in Distributed training: multi-node jobs and the distributors API

    3. 3.Pushing an updated image under the same tag automatically updates an existing CRE.Why is that wrong?

      Snowflake checks the image's SHA digest at runtime and rejects the submission if the image changed after registration. You have to recreate the CRE.

      Covered in Custom runtime images for libraries the base image lacks

    Practise it for real

    Package a pinned XGBoost version into a Custom Runtime Environment and run an ML Job on it

    1. 1.Run snow spcs image-registry login, then pull the cpu_x86_64:2.7.2 base image with docker pull.

      Why: Every custom image has to be built on an official Snowflake ML base image, and Docker needs to be logged in to the registry first.

      You should see: The base image is in your local Docker image list.

    2. 2.Write the Dockerfile from this page and run docker build --platform linux/amd64 -t my-custom-image:v1 .

      Why: The image has to be built for linux/amd64 to run in Snowflake's execution environment.

      You should see: The local image my-custom-image:v1 exists.

    3. 3.Run snow custom-image validate my-custom-image:v1 and fix anything it reports.

      Why: Checking locally catches entrypoint, environment variable and dependency problems before the server-side registration check.

      You should see: The validation report passes.

    4. 4.Tag and push the image to <registry_url>/<database>/<schema>/<image_repo>/my-custom-image:v1, then run CREATE CUSTOM RUNTIME ENVIRONMENT my_cre.

      Why: Registering the image runs server-side validation and creates an object you can grant USAGE on.

      You should see: my_cre is in the RESOLVED state.

    5. 5.Decorate a function with @jobs.remote("MY_COMPUTE_POOL", stage_name="payload_stage", runtime_environment="cre@my_cre") and call it.

      Why: The cre@<name> format points the ML Job at your registered environment.

      You should see: An MLJob object is returned and its status eventually reaches DONE.

    Stuck? Get a nudge

    If the job is rejected, check whether the image tag was pushed again after you created the CRE.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Invoking a @remote decorated function returns a Snowflake MLJob object that can be used to manage and monitor the job execution.”
      ↩︎ Operating ML Jobs
      “submit_from_stage: For running Python projects saved on a Snowflake stage”
      ↩︎ Operating ML Jobs
      “For file-based jobs, use the special __return__ variable to specify the return value.”
      ↩︎ Operating ML Jobs
      “Snowflake automatically uses the latest available version of the Snowflake Container Runtime on your compute pool.”
      ↩︎ Exam trap 1
    2. 2.
      “In multi-node jobs, you can access logs from specific instances:”
      ↩︎ Operating ML Jobs
      “ML Jobs automatically wait for the specified target_instances to be available before executing your payload.”
      ↩︎ Distributed training: multi-node jobs and the distributors API
      “You must set MAX_NODES to be greater than or equal to the number of target instances”
      ↩︎ Exam trap 2
      “the job payload is executed as soon as the minimum number of nodes becomes available”
      ↩︎ Checkpoint
    3. 4.
      “Train models on datasets that are larger than the memory of a single compute node”
      ↩︎ Distributed training: multi-node jobs and the distributors API
      “However, Snowflake defaults to using all available resources.”
      ↩︎ Distributed training: multi-node jobs and the distributors API
    4. 5.
      “Snowpark-optimized warehouses make it possible to use Snowpark stored procedures to run single-node ML training workloads directly in Snowflake.”
      ↩︎ Distributed training: multi-node jobs and the distributors API
    5. 6.
      “All custom images must be derived from an official Snowflake ML runtime base image.”
      ↩︎ Custom runtime images for libraries the base image lacks
      “During this step, Snowflake performs a server-side validation check to confirm the image executes correctly.”
      ↩︎ Custom runtime images for libraries the base image lacks
      “The active role must hold the USAGE privilege on the target Custom Runtime Environment, along with access to the underlying image repository.”
      ↩︎ Custom runtime images for libraries the base image lacks
      “If the underlying image in the repository was modified after the CRE was registered, the job submission is rejected to ensure reproducibility.”
      ↩︎ Exam trap 3
      “Environments marked as DEPRECATED can’t be used for new jobs.”
      ↩︎ Checkpoint
      “Before pushing the image to Snowflake, validate it locally to ensure it meets runtime requirements.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 29 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.