CertSafari
    Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01)· Lessons

    Domain 3 · Lesson 10/17

    Real-time model inference on SPCS: deploy, call and monitor REST endpoints

    Implement inference deployment patterns.

    13 min read
    6% of exam
    4 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Choose between real-time, native batch (SQL) and job-based batch inference for a given workload
    • Deploy a registered model version as an HTTPS service with create_service and know its limits and server defaults
    • Find and call a service's public and internal endpoints, including method-name URL rules
    • Configure instance scaling and track real-time inference metrics from the event table and Snowsight
    • Run job-based batch inference over unstructured files and multimodal models by passing stage paths to run_batch

    Key concept

    One registry, two compute engines — The Snowflake Model Registry can serve a model version from a virtual warehouse or from Snowpark Container Services (SPCS). You choose real-time REST, native batch through SQL, or job-based batch according to how fast you need answers, what form the data takes (HTTP payload, tables or staged files) and how the workload has to scale.

    1.Three inference patterns on two engines

    Snowflake runs inference on two engines: the warehouse, which is the SQL engine, and Snowpark Container Services. The registry gives you three ways to use them. Real-time Inference (REST API) runs on SPCS and answers individual HTTP requests with low latency. It is the pattern for powering external applications. Snowflake Native Batch Inference (SQL) calls the model as a function inside SQL pipelines such as Dynamic Tables, Snowpark, dbt and tasks. Job-based Batch Inference treats inference as a separate compute stage on SPCS and is built for high throughput, including files read directly from stages.

    The pattern follows from the workload. A user-facing app that needs an immediate answer to a small request is a real-time case. Scoring rows that already sit in tables as part of a pipeline is native batch. Scoring millions of images or audio files, or running a historical backfill, is job-based batch.

    Matching the workload to the inference pattern
    PatternEngineData sourceBest for
    Real-time Inference (REST API)Snowpark Container ServicesSmall inputs passed via HTTP payloadWeb/Mobile app backends, high-concurrency request spikes
    Snowflake Native Batch Inference (SQL)Warehouse (SQL engine)Data residing in Snowflake TablesDynamic Tables, Snowpark, dbt, SQL-first users
    Job-based Batch InferenceSnowpark Container ServicesData residing in Snowflake Stages (Files)Images, video, audio; large-scale historical backfills

    The patterns share one cost view. Whether a workload runs on a warehouse or on SPCS, you can query the MODEL_SERVING_USAGE_HISTORY view to see the estimated credits it consumed.

    Checkpoint 1 of 7· Check yourself

    A team needs to score several million video files held in a Snowflake stage, and nobody is waiting on an individual result. Which pattern does Snowflake position for this?

    Sources1

    2.Deploying a model version as an HTTPS service

    Real-time serving packages a registered model as an HTTP server inside SPCS. You don't manage Docker images or Kubernetes clusters, and you get autoscaling, observability, traffic splitting and shadow/canary upgrades. Online inference fits when the app needs low latency, the model backs a web or mobile app, the input fits in an HTTP payload, and the service has to scale horizontally.

    Prerequisites: snowflake-ml-python 1.8.0 or later, a model logged in the registry, and these privileges: USAGE or OWNERSHIP on the compute pool (or use the default System Compute Pools), BIND SERVICE ENDPOINT on the account to create a public endpoint, and OWNER or READ on the model.

    Deployment starts from a model version object. You get one by logging a new version or by fetching an existing one, and then you call create_service:

    Creating a model service from a registered model versionpython
    # reg is a snowflake.ml.registry.Registry object
    example_mv_object = reg.get_model("mymodel_name").version("version_name") # a snowflake.ml.model.ModelVersion object
    
    example_mv_object.create_service(service_name="myservice",
                      service_compute_pool="my_compute_pool",
                      ingress_enabled=True,
                      gpu_requests=None)

    service_name must be unique in the account. service_compute_pool must already exist, and it can be SYSTEM_COMPUTE_POOL_CPU or SYSTEM_COMPUTE_POOL_GPU. ingress_enabled must be True if anything outside Snowflake will call the service. gpu_requests decides whether a model that can run on either CPU or GPU runs on GPUs. If you request GPUs for a CPU-only model type such as scikit-learn, the image build fails. A new service can take up to 10 minutes to create for CPU models and 20 minutes for GPU models. The image is built on the serving pool by default. You can pass image_build_compute_pool to build on a smaller pool, and calling create_service again does not rebuild every time.

    Checkpoint 2 of 7· Fill the gap

    Which argument has to be True so that applications outside Snowflake can call the service?

    example_mv_object.create_service(service_name="myservice",
                      service_compute_pool="my_compute_pool",
                       ? =True,
                      gpu_requests=None)

    Limitations. A model that has a table function can't be deployed. Models built with Snowpark ML modeling classes can't go to GPU environments, so extract the native model and deploy that.

    Server defaults. A CPU model runs (2 × CPUs) + 1 worker processes. A GPU model runs one. You can override this with num_workers. Some models are not thread-safe, so every worker loads its own copy of the model, and with a large model those copies can use up the node's memory. Each instance requests the whole node unless you set cpu_requests, memory_requests or gpu_requests. To scale GPU models, pick the smallest GPU node the model fits in, use gpu_requests=1, and raise max_instances. The endpoint is always named inference on port 5000, and you can't change either.

    Checkpoint 3 of 7· Check yourself

    A large CPU model service keeps running out of memory on its node. Which default is the most likely cause?

    Sources2

    3.Finding and calling the endpoint

    Every service has an internal DNS name. A service deployed with ingress_enabled also gets a public HTTP endpoint, and you can call it through either one.

    - Public: run SHOW ENDPOINTS. The ingress_url column holds a value like unique-service-id-account-id.snowflakecomputing.app. Private link users read privatelink_ingress_url instead. - Internal: DESCRIBE SERVICE returns dns_name, and SHOW ENDPOINTS IN SERVICE returns the port. Inside Snowflake you call http://dns_name:port.

    To call a particular model method, add the method name as the URL path. Underscores in the method name become dashes in the URL. From Python, list_services() returns both endpoints:

    Listing a model version's services and their endpointspython
    # mv: snowflake.ml.model.ModelVersion
    mv.list_services()

    The output includes the public endpoint in inference_endpoint and the internal one in internal_endpoint. You can also manage deployed services from the Model Registry UI in Snowsight.

    Checkpoint 4 of 7· Check yourself

    A web app needs to call the model's predict_proba method through the public ingress URL. Which path does it use?

    Sources2

    4.Scaling the service and tracking its metrics

    Since snowflake-ml-python 1.25.0, create_service accepts min_instances and max_instances. The service starts with min_instances and scales within that range as traffic and hardware use change. Scaling triggers usually fire after the condition has held for about a minute, and the new instances then still have to be provisioned. With the default min_instances=0, the service suspends after 30 minutes without traffic, and the next request has to wait for it to resume. For production, set min_instances to 1 or more.

    Model serving services write performance and health metrics to the event table, covering resource utilization, request rates and latency. You can read them in three ways:

    1. The built-in helper SPCS_GET_METRICS(), which also works for suspended services. 2. A direct event table query that filters on RECORD_TYPE = 'METRIC', the service name and the container. 3. Snowsight: Monitoring » Services & jobs, then pick the service. The page has Logs, Metrics and Events tabs that you can filter by instance and container.

    The inference container is named model-inference. For image build problems, look at model-build. Logs come from SPCS_GET_LOGS(), which requires at least MONITOR on the service, or from SYSTEM$GET_SERVICE_LOGS for live debugging.

    Reading a model service's metrics with the built-in helpersql
    -- Retrieve metrics using the service helper function
    SELECT *
    FROM TABLE(mydb.myschema.my_model_service!SPCS_GET_METRICS())
    WHERE
    timestamp > dateadd(hour, -1, current_timestamp())
    AND instance_id = 0  -- choose all instances or one particular
    AND container_name = 'model-inference';

    Checkpoint 5 of 7· Match them up

    Match each tool to what it gives you for a model service

    Tap a term, then the definition that fits it.

    Checkpoint 6 of 7· Exam question

    A retail mobile app must obtain fraud scores from a registered model in well under a second per call, from an external HTTPS client, with traffic that scales up and down through the day. Which deployment pattern fits this requirement?

    Sources3

    5.Batch inference over unstructured and multimodal data

    Real-time endpoints take small inputs in an HTTP payload. For files, use job-based batch inference: call ModelVersion.run_batch (snowflake-ml-python 2.0.0 or later) on a registered model version. The job builds an inference image, runs on the SPCS compute pool you name without you creating a service, writes results to a stage, and then winds the compute down. Snowflake lists processing images, audio or video files with multimodal models as a use case for run_batch. In SQL, EXECUTE INFERENCE JOB SERVICE runs the same kind of job.

    Passing files. Put fully qualified stage paths in the input DataFrame, and the job reads each file and passes its content to the model. list_stage_files can build that DataFrame from a stage path, optionally filtered by a pattern such as .*\.jpg. Use InputSpec(column_handling=...) to say which column holds stage paths (FULL_STAGE_PATH) and which encoding the model expects: RAW_BYTES, BASE64 or BASE64_DATA_URL.

    Multimodal models. A Hugging Face pipeline such as an image-text-to-text model can be logged for SPCS and run with the vLLM engine through InferenceSpec(engine_options=EngineOptions(...)). Its chat messages can reference image, video or audio files by stage path, and the job downloads and converts them.

    Limits. The output stage must be an internal stage. Input can come from internal stages, or from Amazon S3 external stages that use server-side encryption; Azure Blob Storage and Google Cloud Storage aren't supported. For multimodal use cases, only server-side encryption is supported.

    Checkpoint 7 of 7· Check yourself

    You run run_batch over image files in a stage. How should the input DataFrame identify the files?

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A Snowpark ML modeling-class model can go straight to a GPU compute pool as long as gpu_requests is set.Why is that wrong?

      Snowpark ML modeling classes can't be deployed to GPU environments at all. The workaround is to extract the native model and deploy that.

      Covered in Deploying a model version as an HTTPS service

    2. 2.A model service stays warm by default, so the first request after a quiet period is always fast.Why is that wrong?

      min_instances defaults to 0, which lets the service auto-suspend after 30 minutes without traffic. The next request has to wait for a resume. Set min_instances to 1 or more in production.

      Covered in Scaling the service and tracking its metrics

    3. 3.You can rename the inference endpoint or move it to another port when you create the service.Why is that wrong?

      The endpoint is fixed: it is named inference and listens on port 5000.

      Covered in Deploying a model version as an HTTPS service

    4. 4.Batch inference over staged files can read from Azure Blob Storage or Google Cloud Storage external stages.Why is that wrong?

      For external stages only Amazon S3 with server-side encryption is supported. The output stage must be an internal stage.

      Covered in Batch inference over unstructured and multimodal data

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Designed for low-latency and real-time use cases. Requests are facilitated via HTTP endpoints and are ideal for powering external applications.”
      ↩︎ Three inference patterns on two engines
      “query the MODEL_SERVING_USAGE_HISTORY view”
      ↩︎ Three inference patterns on two engines
      “The Snowflake Model Registry provides a unified interface to both engines.”
      ↩︎ Key concept
      “This is ideal for processing files, such as images, video, and audio, directly from Snowflake Stages.”
      ↩︎ Checkpoint
    2. 2.
      “BIND SERVICE ENDPOINT privilege on account to be able to create a public endpoint.”
      ↩︎ Deploying a model version as an HTTPS service
      “it can take up to 10 minutes to create the service for CPU-powered models and 20 minutes for GPU-powered models.”
      ↩︎ Deploying a model version as an HTTPS service
      “The output contains an ingress_url column, which has an entry of the format unique-service-id-account-id.snowflakecomputing.app.”
      ↩︎ Finding and calling the endpoint
      “Models developed using Snowpark ML modeling classes can’t be deployed to environments that have a GPU.”
      ↩︎ Exam trap 1
      “The inference endpoint is named inference and uses port 5000. These cannot be customized.”
      ↩︎ Exam trap 3
      “Models developed using Snowpark ML modeling classes can’t be deployed to environments that have a GPU.”
      ↩︎ Prediction
      “the service loads a separate copy of the model for each worker process. This can result in resource depletion for large models.”
      ↩︎ Checkpoint
      “the method name predict_proba is changed to predict-proba in the URL.”
      ↩︎ Checkpoint
    3. 3.
      “help you monitor resource utilization, request rates, latency, and other operational characteristics.”
      ↩︎ Scaling the service and tracking its metrics
      “Scaling triggers typically activate after one minute of meeting the required condition.”
      ↩︎ Scaling the service and tracking its metrics
      “If min_instances is set to 0 (the default), the service will automatically suspend if no traffic is detected for 30 minutes.”
      ↩︎ Exam trap 2
      “Model serving services include a built-in helper function that retrieves metrics from the event table for running or suspended services”
      ↩︎ Checkpoint
    4. 4.
      “Process images, audio, or video files using multimodal models with unstructured data.”
      ↩︎ Batch inference over unstructured and multimodal data
      “For multimodal use cases, only server-side encryption is supported.”
      ↩︎ Batch inference over unstructured and multimodal data
      “External stages: Amazon S3 only, and the stage must use server-side encryption.”
      ↩︎ Exam trap 4
      “For unstructured data, reference the files by their fully qualified stage paths in the input DataFrame.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Batch inference in Snowflake: warehouse SQL, service functions and run_batch jobs

    Spotted a mistake, or was something unclear? Tell us.