What you will be able to do
- Run registry-backed batch inference on a virtual warehouse from SQL or Python
- Reuse an online SPCS model service as a SQL service function for batch scoring
- Launch, monitor and read results from a job-based batch inference run with run_batch, with no service involved
- Configure batch jobs over unstructured files for multimodal models within the documented stage constraints
1.Registry-backed batch inference on a warehouse
Native batch inference treats a model version in the Snowflake Model Registry like a SQL function. You can call it from SQL, Snowpark Python, streaming pipelines and Dynamic Tables, and tools such as dbt can use its results. To make a model runnable on a warehouse, include WAREHOUSE in target_platforms when you log it. To check an existing version, run SHOW VERSIONS IN MODEL and look for WAREHOUSE in the runnable_in column.
In SQL you call a method through MODEL(...)!method. Leaving out the version uses the default version, and naming a version or an alias such as LAST pins it:
SELECT MODEL(<model_name>,<version_or_alias_name>)!<method_name>(...) FROM <table_name>;From Python, call run on the model version. The output has the same type as the input: a pandas DataFrame in gives a pandas DataFrame out. Snowpark DataFrames are evaluated lazily, so nothing executes until collect, show or to_pandas. mv.show_functions() and SHOW FUNCTION IN MODEL list the callable methods. If you put the MODEL(...)!predict call inside a Dynamic Table with REFRESH_MODE = INCREMENTAL, only newly arrived rows get scored on each refresh.
A warehouse suits CPU-runnable models such as scikit-learn or XGBoost, and it needs no compute pool. It is the wrong choice when the model needs a GPU or is larger than about 15GB. That limit is lower on smaller warehouses. Large models, GPU models and custom environments belong on SPCS instead.
# Run inference on a warehouse
# mv: snowflake.ml.model.ModelVersion
remote_prediction = mv.run(input_features, function_name="predict")
remote_prediction.show()Checkpoint 1 of 8· Check yourself
A 20GB deep learning model that needs a GPU must score rows in a Snowflake table from SQL. Where should it be hosted?
Warehouses are the wrong runtime for GPU models or models over about 15GB. SPCS is the runtime for large and GPU models.
“Memory Limits: Model size exceeds 15GB. (this limit is lower for smaller warehouse sizes).”Source: docs.snowflake.com
Sources1
2.Batch scoring through the online service as a service function
A model deployed to SPCS as a service (the same kind of service created with create_service for real-time REST calls) can also run SQL batch inference. A warehouse model is called through MODEL(...)!method. A service is called by its own name, service_name!method_name(...), and parameters work the same way: either all positional or all named.
-- Positional
SELECT my_service!predict(input_text, 0.9, 512) FROM my_table;In Python you use the same run call as on a warehouse plus a service_name argument. Three operational rules apply:
- Access. By default only the service owner can call service functions. To let another role call it, grant that role the service role INFERENCE_SERVICE_FUNCTION_USAGE with GRANT SERVICE ROLE.
- Warehouse size. The warehouse only sends requests to the compute pool. A large warehouse runs many threads and can overwhelm the service, so use XSMALL or SMALL.
- Suspension. A suspended service resumes when a request arrives. If the compute pool has no free nodes, the resume can take too long and the query fails. Resume the service explicitly and wait until it is available before running the batch.
Checkpoint 2 of 8· Fill the gap
Which argument sends this run call to the SPCS service instead of the warehouse?
# Run inference on SPCS
# mv: snowflake.ml.model.ModelVersion
remote_prediction = mv.run(input_features, function_name="predict", ? ="example_spcs_service")
remote_prediction.show()Adding service_name to run sends inference to the deployed SPCS service. compute_pool belongs to run_batch, and service_compute_pool to create_service.
Source: docs.snowflake.comCheckpoint 3 of 8· Check yourself
A team routes a nightly SQL batch through a model service on SPCS and the service is getting overwhelmed. What does Snowflake recommend for the calling warehouse?
A warehouse runs many threads per node and sends many requests at once, so a small warehouse keeps the SPCS service from being flooded.
“the recommendation is to utilize an XSMALL or SMALL warehouse when the model is deployed to SPCS.”Source: docs.snowflake.com
Checkpoint 4 of 8· Exam question
A team deployed a model service with a public endpoint and calls it with a Programmatic Access Token. Every request returns HTTP 404, even though the service shows as RUNNING. What is the most likely class of cause?
Correct answer: D — An authorization problem such as an invalid token or missing network route, because all authorization failures surface as 404.
- A. Incorrect: a signature or payload problem produces an error from the model server about the input, not a 404 for every call.
- B. Incorrect: a service reported as RUNNING is on an active pool; pool suspension is not signaled by an endpoint-wide 404.
- C. Incorrect: both dataframe_split and dataframe_records are accepted request layouts, so using one of them does not cause a 404.
- D. Correct: the documentation states that authorization failures, including an incorrect token or lack of a network route to the service, are reported as 404.
Sources1
3.Job-based batch inference with run_batch, with no service
For large batch workloads you don't have to keep a service running. ModelVersion.run_batch (snowflake-ml-python 2.0.0 or later) runs inference as an SPCS job. In SQL the equivalent is EXECUTE INFERENCE JOB SERVICE. Snowflake runs the job in this order:
Checkpoint 5 of 8· Put it in order
Put the stages of a run_batch job in order
- 1.Run inference
- 2.Write results to a stage
- 3.Wind the compute down
- 4.Build the inference image
- 5.Provision an SPCS job on the named compute pool
The job builds its image, provisions compute, runs inference, writes output to a stage and then releases the compute, so you don't pay for idle capacity.
“The job builds an inference image, provisions an SPCS job on the compute pool you name, runs inference, writes results to a stage”Source: docs.snowflake.com
Use it for millions or billions of rows, for inference that runs as its own asynchronous pipeline stage, or as a step in an Airflow DAG or a Snowflake task.
from snowflake.ml.model.batch_inference import OutputSpec, SaveMode
input_df = session.table("my_db.my_schema.my_input_table")
job = mv.run_batch(
input_df,
compute_pool="my_compute_pool",
output_spec=OutputSpec(
stage_location="@my_db.my_schema.my_stage/predictions/",
mode=SaveMode.ERROR,
),
)
job.wait() # optional: block until the job finishesrun_batch is asynchronous by default (async_=True). Pass async_=False if the call should block. Give it exactly one input: a Snowpark DataFrame X (which gets materialized as Parquet under the output path) or an input_stage_location that already holds Parquet. Passing both, or neither, raises a ValueError. The input path must not sit inside the output path.
The call returns an MLJob handle, and the ML Job APIs manage it: job.get_logs(), job.cancel(), list_jobs(), get_job(), delete_job(). The result function is not supported for batch inference jobs, so read predictions from the output stage. A run that has not finished can leave output that looks truncated. Wait for the _SUCCESS file before reading. If a job fails because files already exist under SaveMode.ERROR, point it at a new directory or use SaveMode.OVERWRITE.
Checkpoint 6 of 8· Check yourself
A pipeline step calls job.wait() on a run_batch job and then needs the predictions. How should it get them?
result() isn't supported for batch inference jobs, so output is read from the stage. The _SUCCESS marker shows the job finished.
“The result function in the ML Job APIs isn’t supported for batch inference jobs. Read the job’s output from the output stage instead.”Source: docs.snowflake.com
Sources2
4.Batch jobs over unstructured files and multimodal models
For images, audio or video, the input DataFrame holds fully qualified stage paths rather than feature values. The job reads each file and passes its contents to the model. list_stage_files(session, "@db.schema.stage/path") builds that DataFrame for you and can filter with a regex pattern and set a column_name. To tell the job which column holds the paths and how to encode the file contents, use InputSpec.column_handling. The model can receive FileEncoding.RAW_BYTES, FileEncoding.BASE64 or FileEncoding.BASE64_DATA_URL.
job = mv.run_batch(
X,
compute_pool="my_compute_pool",
output_spec=OutputSpec(stage_location="@my_db.my_schema.my_stage/predictions/"),
input_spec=InputSpec(
# FULL_STAGE_PATH: the column holds a fully qualified path (@db.schema.stage/path) to a file
# RAW_BYTES: download the file and hand its bytes to the model
column_handling={
"path": ColumnHandlingOptions(
input_format=InputFormat.FULL_STAGE_PATH,
convert_to=FileEncoding.RAW_BYTES,
)
}
),
)| Location | What is supported |
|---|---|
| Input: internal stages | All types of internal stages |
| Input: external stages | Amazon S3 only, with server-side encryption; Azure Blob Storage and Google Cloud Storage aren't supported |
| Input: mixed | One DataFrame can reference several stages, internal and external |
| Output: OutputSpec(stage_location=...) | Must be an internal stage |
External S3 input needs a one-time admin setup (a storage integration and IAM permissions on the bucket), and the role running the job needs USAGE on the external stage. Multimodal use cases support only server-side encryption. partition_column doesn't work with Hugging Face pipeline models or FUNCTION-type methods.
Checkpoint 7 of 8· Check yourself
A company will score audio files in an S3 external stage with a registered speech model and keep the predictions for later analysis. Which design is valid?
S3 is the only external input that's supported, and it must use server-side encryption. The output location must be an internal stage.
“The output stage specified by OutputSpec(stage_location=...) must be an internal stage.”Source: docs.snowflake.com
Checkpoint 8 of 8· Exam question
A model exposes a method named `predict_proba`. Which URL path segment should a REST client use to call that method on the service endpoint?
Correct answer: B — `/predict-proba`, because underscores in the method name are replaced by dashes in the endpoint URL.
- A. Incorrect: the underscore form is not used in the URL; the endpoint converts underscores to dashes.
- B. Correct: Snowflake maps underscores in method names to dashes, so predict_proba is served at predict-proba.
- C. Incorrect: the underscores are replaced rather than removed, so the joined name would not match any method.
- D. Incorrect: underscores are not treated as path separators, so a nested path would not route to the method.
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A finished run_batch job hands back its predictions through job.result(), as other ML jobs do.Why is that wrong?
result is not supported for batch inference jobs. Predictions have to be read from the output stage.
Covered in Job-based batch inference with run_batch, with no service
2.Any external stage (S3, Azure or GCS) can feed a batch inference job.Why is that wrong?
The only supported external input is Amazon S3 with server-side encryption. Azure Blob Storage and Google Cloud Storage are not supported.
Covered in Batch jobs over unstructured files and multimodal models
3.Any role with READ on the model can call the deployed service's SQL functions.Why is that wrong?
By default only the service owner can call service functions. Other roles need the INFERENCE_SERVICE_FUNCTION_USAGE service role, granted with GRANT SERVICE ROLE.
Covered in Batch scoring through the online service as a service function
Practise it for real
Run a job-based batch inference over a table with no model service, and read its results safely.
1.Get a ModelVersion with Registry(session=session, database_name=..., schema_name=...).get_model("my_model").version("my_version").
Why: run_batch is a method on the model version object.
You should see: A ModelVersion handle with no service created.
2.Call mv.run_batch(input_df, compute_pool="my_compute_pool", output_spec=OutputSpec(stage_location="@my_db.my_schema.my_stage/predictions/", mode=SaveMode.ERROR)).
Why: SaveMode.ERROR stops the job from mixing new output with files already in the directory.
You should see: The call returns an MLJob handle at once, because async_ defaults to True.
3.Call job.wait(), and if the job fails, print(job.get_logs()).
Why: The ML Job APIs are how you inspect and troubleshoot batch jobs.
You should see: The job completes, or the logs show why it failed.
4.List the output stage path and check for the _SUCCESS file before you read predictions from it.
Why: Output from an unfinished job can look truncated, and result() is not supported.
You should see: A _SUCCESS file next to the prediction files.
Stuck? Get a nudge
If the job fails because output files already exist, use a new directory or SaveMode.OVERWRITE.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.snowflake.com/en/developer-guide/snowflake-ml/inference/native-batch-inference-sqlOfficial docs
“To host your model on a warehouse, specify WAREHOUSE in the target_platforms argument to log your model.”
↩︎ Registry-backed batch inference on a warehouse“If the runnable_in column has WAREHOUSE as a value, you can run it.”
↩︎ Registry-backed batch inference on a warehouse“Unlike running in a warehouse, functions can be called from a service by calling service_name!method_name(...).”
↩︎ Batch scoring through the online service as a service function“To mitigate this risk, you can explicitly resume the service and wait for its availability.”
↩︎ Batch scoring through the online service as a service function“By default, only the service owner can invoke service functions.”
↩︎ Exam trap 3“Memory Limits: Model size exceeds 15GB. (this limit is lower for smaller warehouse sizes).”
↩︎ Checkpoint“the recommendation is to utilize an XSMALL or SMALL warehouse when the model is deployed to SPCS.”
↩︎ Checkpoint - 2.https://docs.snowflake.com/en/developer-guide/snowflake-ml/inference/batch-inference-jobsOfficial docs
“Integrate inference as a step within an Airflow DAG or a Snowflake task.”
↩︎ Job-based batch inference with run_batch, with no service“Check for the _SUCCESS file before reading results.”
↩︎ Job-based batch inference with run_batch, with no service“For unstructured data, reference the files by their fully qualified stage paths in the input DataFrame.”
↩︎ Batch jobs over unstructured files and multimodal models“For multimodal use cases, only server-side encryption is supported.”
↩︎ Batch jobs over unstructured files and multimodal models“The result function in the ML Job APIs isn’t supported for batch inference jobs.”
↩︎ Exam trap 1“External stages: Amazon S3 only, and the stage must use server-side encryption.”
↩︎ Exam trap 2“The job builds an inference image, provisions an SPCS job on the compute pool you name, runs inference, writes results to a stage”
↩︎ Checkpoint“The result function in the ML Job APIs isn’t supported for batch inference jobs. Read the job’s output from the output stage instead.”
↩︎ Checkpoint“The output stage specified by OutputSpec(stage_location=...) must be an internal stage.”
↩︎ Checkpoint