CertSafari
    Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01)· Lessons

    Domain 3 · Lesson 10/17

    Batch inference in Snowflake: warehouse SQL, service functions and run_batch jobs

    Implement inference deployment patterns.

    12 min read
    6% of exam
    2 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Run registry-backed batch inference on a virtual warehouse from SQL or Python
    • Reuse an online SPCS model service as a SQL service function for batch scoring
    • Launch, monitor and read results from a job-based batch inference run with run_batch, with no service involved
    • Configure batch jobs over unstructured files for multimodal models within the documented stage constraints

    1.Registry-backed batch inference on a warehouse

    Native batch inference treats a model version in the Snowflake Model Registry like a SQL function. You can call it from SQL, Snowpark Python, streaming pipelines and Dynamic Tables, and tools such as dbt can use its results. To make a model runnable on a warehouse, include WAREHOUSE in target_platforms when you log it. To check an existing version, run SHOW VERSIONS IN MODEL and look for WAREHOUSE in the runnable_in column.

    In SQL you call a method through MODEL(...)!method. Leaving out the version uses the default version, and naming a version or an alias such as LAST pins it:

    Calling a method of a specific model version or alias from SQLsql
    SELECT MODEL(<model_name>,<version_or_alias_name>)!<method_name>(...) FROM <table_name>;

    From Python, call run on the model version. The output has the same type as the input: a pandas DataFrame in gives a pandas DataFrame out. Snowpark DataFrames are evaluated lazily, so nothing executes until collect, show or to_pandas. mv.show_functions() and SHOW FUNCTION IN MODEL list the callable methods. If you put the MODEL(...)!predict call inside a Dynamic Table with REFRESH_MODE = INCREMENTAL, only newly arrived rows get scored on each refresh.

    A warehouse suits CPU-runnable models such as scikit-learn or XGBoost, and it needs no compute pool. It is the wrong choice when the model needs a GPU or is larger than about 15GB. That limit is lower on smaller warehouses. Large models, GPU models and custom environments belong on SPCS instead.

    Running registry-backed inference on a warehouse from Pythonpython
    # Run inference on a warehouse
    # mv: snowflake.ml.model.ModelVersion
    remote_prediction = mv.run(input_features, function_name="predict")
    remote_prediction.show()

    Checkpoint 1 of 8· Check yourself

    A 20GB deep learning model that needs a GPU must score rows in a Snowflake table from SQL. Where should it be hosted?

    Sources1

    2.Batch scoring through the online service as a service function

    A model deployed to SPCS as a service (the same kind of service created with create_service for real-time REST calls) can also run SQL batch inference. A warehouse model is called through MODEL(...)!method. A service is called by its own name, service_name!method_name(...), and parameters work the same way: either all positional or all named.

    Calling a deployed model service as a SQL functionsql
    -- Positional
    SELECT my_service!predict(input_text, 0.9, 512) FROM my_table;

    In Python you use the same run call as on a warehouse plus a service_name argument. Three operational rules apply:

    - Access. By default only the service owner can call service functions. To let another role call it, grant that role the service role INFERENCE_SERVICE_FUNCTION_USAGE with GRANT SERVICE ROLE. - Warehouse size. The warehouse only sends requests to the compute pool. A large warehouse runs many threads and can overwhelm the service, so use XSMALL or SMALL. - Suspension. A suspended service resumes when a request arrives. If the compute pool has no free nodes, the resume can take too long and the query fails. Resume the service explicitly and wait until it is available before running the batch.

    Checkpoint 2 of 8· Fill the gap

    Which argument sends this run call to the SPCS service instead of the warehouse?

    # Run inference on SPCS
    # mv: snowflake.ml.model.ModelVersion
    remote_prediction = mv.run(input_features, function_name="predict",  ? ="example_spcs_service")
    remote_prediction.show()

    Checkpoint 3 of 8· Check yourself

    A team routes a nightly SQL batch through a model service on SPCS and the service is getting overwhelmed. What does Snowflake recommend for the calling warehouse?

    Checkpoint 4 of 8· Exam question

    A team deployed a model service with a public endpoint and calls it with a Programmatic Access Token. Every request returns HTTP 404, even though the service shows as RUNNING. What is the most likely class of cause?

    Sources1

    3.Job-based batch inference with run_batch, with no service

    For large batch workloads you don't have to keep a service running. ModelVersion.run_batch (snowflake-ml-python 2.0.0 or later) runs inference as an SPCS job. In SQL the equivalent is EXECUTE INFERENCE JOB SERVICE. Snowflake runs the job in this order:

    Checkpoint 5 of 8· Put it in order

    Put the stages of a run_batch job in order

    1. 1.Run inference
    2. 2.Write results to a stage
    3. 3.Wind the compute down
    4. 4.Build the inference image
    5. 5.Provision an SPCS job on the named compute pool

    Use it for millions or billions of rows, for inference that runs as its own asynchronous pipeline stage, or as a step in an Airflow DAG or a Snowflake task.

    Submitting a batch inference job and waiting for itpython
    from snowflake.ml.model.batch_inference import OutputSpec, SaveMode
    
    input_df = session.table("my_db.my_schema.my_input_table")
    
    job = mv.run_batch(
        input_df,
        compute_pool="my_compute_pool",
        output_spec=OutputSpec(
            stage_location="@my_db.my_schema.my_stage/predictions/",
            mode=SaveMode.ERROR,
        ),
    )
    
    job.wait()  # optional: block until the job finishes

    run_batch is asynchronous by default (async_=True). Pass async_=False if the call should block. Give it exactly one input: a Snowpark DataFrame X (which gets materialized as Parquet under the output path) or an input_stage_location that already holds Parquet. Passing both, or neither, raises a ValueError. The input path must not sit inside the output path.

    The call returns an MLJob handle, and the ML Job APIs manage it: job.get_logs(), job.cancel(), list_jobs(), get_job(), delete_job(). The result function is not supported for batch inference jobs, so read predictions from the output stage. A run that has not finished can leave output that looks truncated. Wait for the _SUCCESS file before reading. If a job fails because files already exist under SaveMode.ERROR, point it at a new directory or use SaveMode.OVERWRITE.

    Checkpoint 6 of 8· Check yourself

    A pipeline step calls job.wait() on a run_batch job and then needs the predictions. How should it get them?

    Sources2

    4.Batch jobs over unstructured files and multimodal models

    For images, audio or video, the input DataFrame holds fully qualified stage paths rather than feature values. The job reads each file and passes its contents to the model. list_stage_files(session, "@db.schema.stage/path") builds that DataFrame for you and can filter with a regex pattern and set a column_name. To tell the job which column holds the paths and how to encode the file contents, use InputSpec.column_handling. The model can receive FileEncoding.RAW_BYTES, FileEncoding.BASE64 or FileEncoding.BASE64_DATA_URL.

    Declaring a stage-path column and the encoding the model expectspython
    job = mv.run_batch(
        X,
        compute_pool="my_compute_pool",
        output_spec=OutputSpec(stage_location="@my_db.my_schema.my_stage/predictions/"),
        input_spec=InputSpec(
            # FULL_STAGE_PATH: the column holds a fully qualified path (@db.schema.stage/path) to a file
            # RAW_BYTES: download the file and hand its bytes to the model
            column_handling={
                "path": ColumnHandlingOptions(
                    input_format=InputFormat.FULL_STAGE_PATH,
                    convert_to=FileEncoding.RAW_BYTES,
                )
            }
        ),
    )
    Stage constraints for batch inference jobs
    LocationWhat is supported
    Input: internal stagesAll types of internal stages
    Input: external stagesAmazon S3 only, with server-side encryption; Azure Blob Storage and Google Cloud Storage aren't supported
    Input: mixedOne DataFrame can reference several stages, internal and external
    Output: OutputSpec(stage_location=...)Must be an internal stage

    External S3 input needs a one-time admin setup (a storage integration and IAM permissions on the bucket), and the role running the job needs USAGE on the external stage. Multimodal use cases support only server-side encryption. partition_column doesn't work with Hugging Face pipeline models or FUNCTION-type methods.

    Checkpoint 7 of 8· Check yourself

    A company will score audio files in an S3 external stage with a registered speech model and keep the predictions for later analysis. Which design is valid?

    Checkpoint 8 of 8· Exam question

    A model exposes a method named `predict_proba`. Which URL path segment should a REST client use to call that method on the service endpoint?

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A finished run_batch job hands back its predictions through job.result(), as other ML jobs do.Why is that wrong?

      result is not supported for batch inference jobs. Predictions have to be read from the output stage.

      Covered in Job-based batch inference with run_batch, with no service

    2. 2.Any external stage (S3, Azure or GCS) can feed a batch inference job.Why is that wrong?

      The only supported external input is Amazon S3 with server-side encryption. Azure Blob Storage and Google Cloud Storage are not supported.

      Covered in Batch jobs over unstructured files and multimodal models

    3. 3.Any role with READ on the model can call the deployed service's SQL functions.Why is that wrong?

      By default only the service owner can call service functions. Other roles need the INFERENCE_SERVICE_FUNCTION_USAGE service role, granted with GRANT SERVICE ROLE.

      Covered in Batch scoring through the online service as a service function

    Practise it for real

    Run a job-based batch inference over a table with no model service, and read its results safely.

    1. 1.Get a ModelVersion with Registry(session=session, database_name=..., schema_name=...).get_model("my_model").version("my_version").

      Why: run_batch is a method on the model version object.

      You should see: A ModelVersion handle with no service created.

    2. 2.Call mv.run_batch(input_df, compute_pool="my_compute_pool", output_spec=OutputSpec(stage_location="@my_db.my_schema.my_stage/predictions/", mode=SaveMode.ERROR)).

      Why: SaveMode.ERROR stops the job from mixing new output with files already in the directory.

      You should see: The call returns an MLJob handle at once, because async_ defaults to True.

    3. 3.Call job.wait(), and if the job fails, print(job.get_logs()).

      Why: The ML Job APIs are how you inspect and troubleshoot batch jobs.

      You should see: The job completes, or the logs show why it failed.

    4. 4.List the output stage path and check for the _SUCCESS file before you read predictions from it.

      Why: Output from an unfinished job can look truncated, and result() is not supported.

      You should see: A _SUCCESS file next to the prediction files.

    Stuck? Get a nudge

    If the job fails because output files already exist, use a new directory or SaveMode.OVERWRITE.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “To host your model on a warehouse, specify WAREHOUSE in the target_platforms argument to log your model.”
      ↩︎ Registry-backed batch inference on a warehouse
      “If the runnable_in column has WAREHOUSE as a value, you can run it.”
      ↩︎ Registry-backed batch inference on a warehouse
      “Unlike running in a warehouse, functions can be called from a service by calling service_name!method_name(...).”
      ↩︎ Batch scoring through the online service as a service function
      “To mitigate this risk, you can explicitly resume the service and wait for its availability.”
      ↩︎ Batch scoring through the online service as a service function
      “By default, only the service owner can invoke service functions.”
      ↩︎ Exam trap 3
      “Memory Limits: Model size exceeds 15GB. (this limit is lower for smaller warehouse sizes).”
      ↩︎ Checkpoint
      “the recommendation is to utilize an XSMALL or SMALL warehouse when the model is deployed to SPCS.”
      ↩︎ Checkpoint
    2. 2.
      “Integrate inference as a step within an Airflow DAG or a Snowflake task.”
      ↩︎ Job-based batch inference with run_batch, with no service
      “Check for the _SUCCESS file before reading results.”
      ↩︎ Job-based batch inference with run_batch, with no service
      “For unstructured data, reference the files by their fully qualified stage paths in the input DataFrame.”
      ↩︎ Batch jobs over unstructured files and multimodal models
      “For multimodal use cases, only server-side encryption is supported.”
      ↩︎ Batch jobs over unstructured files and multimodal models
      “The result function in the ML Job APIs isn’t supported for batch inference jobs.”
      ↩︎ Exam trap 1
      “External stages: Amazon S3 only, and the stage must use server-side encryption.”
      ↩︎ Exam trap 2
      “The job builds an inference image, provisions an SPCS job on the compute pool you name, runs inference, writes results to a stage”
      ↩︎ Checkpoint
      “The result function in the ML Job APIs isn’t supported for batch inference jobs. Read the job’s output from the output stage instead.”
      ↩︎ Checkpoint
      “The output stage specified by OutputSpec(stage_location=...) must be an internal stage.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 21 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.