What you will be able to do
- Choose between real-time, native batch (SQL) and job-based batch inference for a given workload
- Deploy a registered model version as an HTTPS service with create_service and know its limits and server defaults
- Find and call a service's public and internal endpoints, including method-name URL rules
- Configure instance scaling and track real-time inference metrics from the event table and Snowsight
- Run job-based batch inference over unstructured files and multimodal models by passing stage paths to run_batch
Key concept
One registry, two compute engines — The Snowflake Model Registry can serve a model version from a virtual warehouse or from Snowpark Container Services (SPCS). You choose real-time REST, native batch through SQL, or job-based batch according to how fast you need answers, what form the data takes (HTTP payload, tables or staged files) and how the workload has to scale.
1.Three inference patterns on two engines
Snowflake runs inference on two engines: the warehouse, which is the SQL engine, and Snowpark Container Services. The registry gives you three ways to use them. Real-time Inference (REST API) runs on SPCS and answers individual HTTP requests with low latency. It is the pattern for powering external applications. Snowflake Native Batch Inference (SQL) calls the model as a function inside SQL pipelines such as Dynamic Tables, Snowpark, dbt and tasks. Job-based Batch Inference treats inference as a separate compute stage on SPCS and is built for high throughput, including files read directly from stages.
The pattern follows from the workload. A user-facing app that needs an immediate answer to a small request is a real-time case. Scoring rows that already sit in tables as part of a pipeline is native batch. Scoring millions of images or audio files, or running a historical backfill, is job-based batch.
| Pattern | Engine | Data source | Best for |
|---|---|---|---|
| Real-time Inference (REST API) | Snowpark Container Services | Small inputs passed via HTTP payload | Web/Mobile app backends, high-concurrency request spikes |
| Snowflake Native Batch Inference (SQL) | Warehouse (SQL engine) | Data residing in Snowflake Tables | Dynamic Tables, Snowpark, dbt, SQL-first users |
| Job-based Batch Inference | Snowpark Container Services | Data residing in Snowflake Stages (Files) | Images, video, audio; large-scale historical backfills |
The patterns share one cost view. Whether a workload runs on a warehouse or on SPCS, you can query the MODEL_SERVING_USAGE_HISTORY view to see the estimated credits it consumed.
Checkpoint 1 of 7· Check yourself
A team needs to score several million video files held in a Snowflake stage, and nobody is waiting on an individual result. Which pattern does Snowflake position for this?
Job-based batch separates inference into its own high-throughput stage and is aimed at files such as images, video and audio held in stages.
“This is ideal for processing files, such as images, video, and audio, directly from Snowflake Stages.”Source: docs.snowflake.com
Sources1
2.Deploying a model version as an HTTPS service
Real-time serving packages a registered model as an HTTP server inside SPCS. You don't manage Docker images or Kubernetes clusters, and you get autoscaling, observability, traffic splitting and shadow/canary upgrades. Online inference fits when the app needs low latency, the model backs a web or mobile app, the input fits in an HTTP payload, and the service has to scale horizontally.
Prerequisites: snowflake-ml-python 1.8.0 or later, a model logged in the registry, and these privileges: USAGE or OWNERSHIP on the compute pool (or use the default System Compute Pools), BIND SERVICE ENDPOINT on the account to create a public endpoint, and OWNER or READ on the model.
Deployment starts from a model version object. You get one by logging a new version or by fetching an existing one, and then you call create_service:
# reg is a snowflake.ml.registry.Registry object
example_mv_object = reg.get_model("mymodel_name").version("version_name") # a snowflake.ml.model.ModelVersion object
example_mv_object.create_service(service_name="myservice",
service_compute_pool="my_compute_pool",
ingress_enabled=True,
gpu_requests=None)service_name must be unique in the account. service_compute_pool must already exist, and it can be SYSTEM_COMPUTE_POOL_CPU or SYSTEM_COMPUTE_POOL_GPU. ingress_enabled must be True if anything outside Snowflake will call the service. gpu_requests decides whether a model that can run on either CPU or GPU runs on GPUs. If you request GPUs for a CPU-only model type such as scikit-learn, the image build fails. A new service can take up to 10 minutes to create for CPU models and 20 minutes for GPU models. The image is built on the serving pool by default. You can pass image_build_compute_pool to build on a smaller pool, and calling create_service again does not rebuild every time.
Checkpoint 2 of 7· Fill the gap
Which argument has to be True so that applications outside Snowflake can call the service?
example_mv_object.create_service(service_name="myservice",
service_compute_pool="my_compute_pool",
? =True,
gpu_requests=None)ingress_enabled=True creates the public HTTP endpoint. BIND SERVICE ENDPOINT is an account privilege you need, not an argument to create_service.
Source: docs.snowflake.comLimitations. A model that has a table function can't be deployed. Models built with Snowpark ML modeling classes can't go to GPU environments, so extract the native model and deploy that.
Server defaults. A CPU model runs (2 × CPUs) + 1 worker processes. A GPU model runs one. You can override this with num_workers. Some models are not thread-safe, so every worker loads its own copy of the model, and with a large model those copies can use up the node's memory. Each instance requests the whole node unless you set cpu_requests, memory_requests or gpu_requests. To scale GPU models, pick the smallest GPU node the model fits in, use gpu_requests=1, and raise max_instances. The endpoint is always named inference on port 5000, and you can't change either.
Checkpoint 3 of 7· Check yourself
A large CPU model service keeps running out of memory on its node. Which default is the most likely cause?
Because models may not be thread-safe, each worker holds a full copy. With many workers and a large model, memory runs out. Lowering num_workers or setting memory_requests helps.
“the service loads a separate copy of the model for each worker process. This can result in resource depletion for large models.”Source: docs.snowflake.com
Sources2
3.Finding and calling the endpoint
Every service has an internal DNS name. A service deployed with ingress_enabled also gets a public HTTP endpoint, and you can call it through either one.
- Public: run SHOW ENDPOINTS. The ingress_url column holds a value like unique-service-id-account-id.snowflakecomputing.app. Private link users read privatelink_ingress_url instead.
- Internal: DESCRIBE SERVICE returns dns_name, and SHOW ENDPOINTS IN SERVICE returns the port. Inside Snowflake you call http://dns_name:port.
To call a particular model method, add the method name as the URL path. Underscores in the method name become dashes in the URL. From Python, list_services() returns both endpoints:
# mv: snowflake.ml.model.ModelVersion
mv.list_services()The output includes the public endpoint in inference_endpoint and the internal one in internal_endpoint. You can also manage deployed services from the Model Registry UI in Snowsight.
Checkpoint 4 of 7· Check yourself
A web app needs to call the model's predict_proba method through the public ingress URL. Which path does it use?
The method name is the URL path, and its underscores are replaced with dashes.
“the method name predict_proba is changed to predict-proba in the URL.”Source: docs.snowflake.com
Sources2
4.Scaling the service and tracking its metrics
Since snowflake-ml-python 1.25.0, create_service accepts min_instances and max_instances. The service starts with min_instances and scales within that range as traffic and hardware use change. Scaling triggers usually fire after the condition has held for about a minute, and the new instances then still have to be provisioned. With the default min_instances=0, the service suspends after 30 minutes without traffic, and the next request has to wait for it to resume. For production, set min_instances to 1 or more.
Model serving services write performance and health metrics to the event table, covering resource utilization, request rates and latency. You can read them in three ways:
1. The built-in helper SPCS_GET_METRICS(), which also works for suspended services.
2. A direct event table query that filters on RECORD_TYPE = 'METRIC', the service name and the container.
3. Snowsight: Monitoring » Services & jobs, then pick the service. The page has Logs, Metrics and Events tabs that you can filter by instance and container.
The inference container is named model-inference. For image build problems, look at model-build. Logs come from SPCS_GET_LOGS(), which requires at least MONITOR on the service, or from SYSTEM$GET_SERVICE_LOGS for live debugging.
-- Retrieve metrics using the service helper function
SELECT *
FROM TABLE(mydb.myschema.my_model_service!SPCS_GET_METRICS())
WHERE
timestamp > dateadd(hour, -1, current_timestamp())
AND instance_id = 0 -- choose all instances or one particular
AND container_name = 'model-inference';Checkpoint 5 of 7· Match them up
Match each tool to what it gives you for a model service
Tap a term, then the definition that fits it.
Metrics and logs each have their own service helper function. SYSTEM$GET_SERVICE_LOGS is for live debugging, and the usage view tracks cost.
“Model serving services include a built-in helper function that retrieves metrics from the event table for running or suspended services”Source: docs.snowflake.com
Checkpoint 6 of 7· Exam question
A retail mobile app must obtain fraud scores from a registered model in well under a second per call, from an external HTTPS client, with traffic that scales up and down through the day. Which deployment pattern fits this requirement?
Correct answer: C — Create a model service on an SPCS compute pool with `ingress_enabled=True` and call its HTTPS endpoint from the backend.
- A. Incorrect: routing every prediction through a warehouse SQL API adds query compilation and warehouse dependency, and does not provide a managed model endpoint.
- B. Incorrect: warehouse SQL inference suits batch scoring integrated with SQL pipelines; per-request statements add query overhead and are not an external HTTPS inference API.
- C. Correct: a model service on SPCS exposes a managed HTTPS REST endpoint designed for low-latency, real-time inference and scales horizontally with its compute pool.
- D. Incorrect: run_batch builds an image and starts an SPCS job per invocation, which has startup latency measured in minutes and is meant for high-volume stages.
Sources3
5.Batch inference over unstructured and multimodal data
Real-time endpoints take small inputs in an HTTP payload. For files, use job-based batch inference: call ModelVersion.run_batch (snowflake-ml-python 2.0.0 or later) on a registered model version. The job builds an inference image, runs on the SPCS compute pool you name without you creating a service, writes results to a stage, and then winds the compute down. Snowflake lists processing images, audio or video files with multimodal models as a use case for run_batch. In SQL, EXECUTE INFERENCE JOB SERVICE runs the same kind of job.
Passing files. Put fully qualified stage paths in the input DataFrame, and the job reads each file and passes its content to the model. list_stage_files can build that DataFrame from a stage path, optionally filtered by a pattern such as .*\.jpg. Use InputSpec(column_handling=...) to say which column holds stage paths (FULL_STAGE_PATH) and which encoding the model expects: RAW_BYTES, BASE64 or BASE64_DATA_URL.
Multimodal models. A Hugging Face pipeline such as an image-text-to-text model can be logged for SPCS and run with the vLLM engine through InferenceSpec(engine_options=EngineOptions(...)). Its chat messages can reference image, video or audio files by stage path, and the job downloads and converts them.
Limits. The output stage must be an internal stage. Input can come from internal stages, or from Amazon S3 external stages that use server-side encryption; Azure Blob Storage and Google Cloud Storage aren't supported. For multimodal use cases, only server-side encryption is supported.
Checkpoint 7 of 7· Check yourself
You run run_batch over image files in a stage. How should the input DataFrame identify the files?
The job reads each file from its full stage path and passes the content to the model, so no service endpoint is involved.
“For unstructured data, reference the files by their fully qualified stage paths in the input DataFrame.”Source: docs.snowflake.com
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A Snowpark ML modeling-class model can go straight to a GPU compute pool as long as gpu_requests is set.Why is that wrong?
Snowpark ML modeling classes can't be deployed to GPU environments at all. The workaround is to extract the native model and deploy that.
2.A model service stays warm by default, so the first request after a quiet period is always fast.Why is that wrong?
min_instances defaults to 0, which lets the service auto-suspend after 30 minutes without traffic. The next request has to wait for a resume. Set min_instances to 1 or more in production.
Covered in Scaling the service and tracking its metrics
3.You can rename the inference endpoint or move it to another port when you create the service.Why is that wrong?
The endpoint is fixed: it is named inference and listens on port 5000.
4.Batch inference over staged files can read from Azure Blob Storage or Google Cloud Storage external stages.Why is that wrong?
For external stages only Amazon S3 with server-side encryption is supported. The output stage must be an internal stage.
Covered in Batch inference over unstructured and multimodal data
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.snowflake.com/en/developer-guide/snowflake-ml/inference/inference-overviewOfficial docs
“Designed for low-latency and real-time use cases. Requests are facilitated via HTTP endpoints and are ideal for powering external applications.”
↩︎ Three inference patterns on two engines“query the MODEL_SERVING_USAGE_HISTORY view”
↩︎ Three inference patterns on two engines“The Snowflake Model Registry provides a unified interface to both engines.”
↩︎ Key concept“This is ideal for processing files, such as images, video, and audio, directly from Snowflake Stages.”
↩︎ Checkpoint - 2.https://docs.snowflake.com/en/developer-guide/snowflake-ml/inference/real-time-inference-rest-apiOfficial docs
“BIND SERVICE ENDPOINT privilege on account to be able to create a public endpoint.”
↩︎ Deploying a model version as an HTTPS service“it can take up to 10 minutes to create the service for CPU-powered models and 20 minutes for GPU-powered models.”
↩︎ Deploying a model version as an HTTPS service“The output contains an ingress_url column, which has an entry of the format unique-service-id-account-id.snowflakecomputing.app.”
↩︎ Finding and calling the endpoint“Models developed using Snowpark ML modeling classes can’t be deployed to environments that have a GPU.”
↩︎ Exam trap 1“The inference endpoint is named inference and uses port 5000. These cannot be customized.”
↩︎ Exam trap 3“Models developed using Snowpark ML modeling classes can’t be deployed to environments that have a GPU.”
↩︎ Prediction“the service loads a separate copy of the model for each worker process. This can result in resource depletion for large models.”
↩︎ Checkpoint“the method name predict_proba is changed to predict-proba in the URL.”
↩︎ Checkpoint - 3.https://docs.snowflake.com/en/developer-guide/snowflake-ml/inference/service-managementOfficial docs
“help you monitor resource utilization, request rates, latency, and other operational characteristics.”
↩︎ Scaling the service and tracking its metrics“Scaling triggers typically activate after one minute of meeting the required condition.”
↩︎ Scaling the service and tracking its metrics“If min_instances is set to 0 (the default), the service will automatically suspend if no traffic is detected for 30 minutes.”
↩︎ Exam trap 2“Model serving services include a built-in helper function that retrieves metrics from the event table for running or suspended services”
↩︎ Checkpoint - 4.https://docs.snowflake.com/en/developer-guide/snowflake-ml/inference/batch-inference-jobsOfficial docs
“Process images, audio, or video files using multimodal models with unstructured data.”
↩︎ Batch inference over unstructured and multimodal data“For multimodal use cases, only server-side encryption is supported.”
↩︎ Batch inference over unstructured and multimodal data“External stages: Amazon S3 only, and the stage must use server-side encryption.”
↩︎ Exam trap 4“For unstructured data, reference the files by their fully qualified stage paths in the input DataFrame.”
↩︎ Checkpoint