What you will be able to do
- Log a trained Python model into a schema with log_model and add new versions
- Retrieve a model version and run inference from Python or with SQL model methods
- Promote a version to production using the default version or an alias
- Choose USAGE or READ for consumers and decide between warehouse and Snowpark Container Services inference
- Name the required create_service arguments for serving a model in Snowpark Container Services
Key concept
Model as a schema-level object — Logging a model in the registry turns a Python model into a Snowflake object that lives in a schema, like a table. Each version exposes methods you call for inference, either on a warehouse or as a service in Snowpark Container Services.
1.Logging a model into the registry
In Snowflake, putting a model into production starts with the Model Registry. Any schema can act as a registry. You don't need to set it up first, but Snowflake recommends a dedicated schema such as ML.REGISTRY. You open the registry with the Registry class, then add models through the reference it returns. The Python API has three main classes: Registry manages the models in a schema, Model represents a model, and ModelVersion represents one version of it.
The registry has built-in support for the most common model types: scikit-learn, xgboost, LightGBM, Prophet, CatBoost, PyTorch, TensorFlow, Keras, Hugging Face pipelines, Sentence Transformer and MLFlow pyfunc. It also accepts other previously trained models and custom processing code. One boundary matters: models trained with Snowflake ML Functions such as FORECAST do not appear in the registry at all.
from snowflake.ml.registry import Registry
reg = Registry(session=session, database_name="ML", schema_name="REGISTRY")Adding a model is called *logging* it. log_model serializes the Python object, which must be pickleable, and creates a Snowflake model object from it. Each call adds one version. To add another version, call log_model again with the same model_name and a new version_name. The name and version together must be unique in the schema. In the same call you can attach a comment, metrics and a sample of the input data, and you can list conda_dependencies so those packages are deployed with the model. You cannot set tags here: tags belong to the model, and the model object only exists once its first version has been logged. You can add tags afterwards.
from snowflake.ml.model import task, type_hints
mv = reg.log_model(clf,
model_name="my_model",
version_name="v1",
conda_dependencies=["scikit-learn"],
comment="My awesome ML model",
metrics={"score": 96},
sample_input_data=train_features,
task=task.Task.TABULAR_BINARY_CLASSIFICATION)Checkpoint 1 of 7· Fill the gap
Which registry method completes this sample and creates the model version?
from snowflake.ml.model import task, type_hints
mv = reg. ? (clf,
model_name="my_model",
version_name="v1",
conda_dependencies=["scikit-learn"],
comment="My awesome ML model",
metrics={"score": 96},
sample_input_data=train_features,
task=task.Task.TABULAR_BINARY_CLASSIFICATION)Adding a model to the registry is called logging, and log_model is the method that serializes the object and creates the model version.
Source: docs.snowflake.comSources1
2.Retrieving a version and running inference
After a model is logged, you call its methods to run inference. Snowflake treats these methods like functions or stored procedures. From Python, you get the ModelVersion back from the registry by model name and version name.
from snowflake.ml.registry import Registry
registry = Registry(session=session, database_name=DATABASE, schema_name=REGISTRY_SCHEMA)
mv = registry.get_model('my_model').version('my_version') # returns ModelVersionThe version's run method runs a batch inference job on either a warehouse or an SPCS compute pool. You pass it a Snowpark or pandas DataFrame, and it returns the same kind: pandas in, pandas out. Snowpark DataFrames are evaluated lazily, so nothing runs until you call collect, show or to_pandas.
You can call the same methods from SQL. my_model!predict(...) calls the default version. MODEL(my_model, <version or alias>)!predict(...) calls a specific version by name or alias.
Two arguments of run are worth knowing. function_name names the method to call, such as predict; mv.show_functions lists the methods a version offers. To run on SPCS instead of a warehouse, add the service_name argument to the same call. In SQL, the default-version form is MODEL(<model_name>)!<method_name>(...), with the table that holds the inference data in the FROM clause. To target a specific version, use MODEL(<model_name>,<version_or_alias_name>)!<method_name>(...). A service deployed in SPCS is called with service_name!method_name(...).
remote_prediction = mv.run(input_features, function_name="predict", service_name="example_spcs_service")Checkpoint 2 of 7· Check yourself
A data scientist calls mv.run() on a model version and passes a pandas DataFrame. What do they get back?
run returns the same kind of DataFrame it received, so pandas input produces pandas output.
“The run method returns a dataframe that matches the type of the dataframe that you’ve specified.”Source: docs.snowflake.com
3.Promoting a version: default version and aliases
The simplest way to manage production is to treat the default version as the production version. Production code calls only the default, and you release a new model by changing which version is the default. When the first version is logged, it becomes the default automatically. Because a model can never be without a default, the docs suggest one workaround if you need to block use before anything is ready: log an initial version that immediately throws an error, and keep it as the default until a real version is ready.
ALTER MODEL my_model SET DEFAULT_VERSION = new_version;If you need more stages than development and production, use aliases. An alias is a label that can sit on only one version at a time, for example alpha, beta or production. Consumers call MODEL(my_model, production)!predict(...). To promote a version, you move the label, and consumer code stays the same. If a role other than the model owner must approve releases, you can put the live version name in a tag instead and read it with SYSTEM$GET_TAG. Another option is to copy approved models into a separate production schema.
SELECT MODEL(my_model, production)!predict(...) FROM ...;Checkpoint 3 of 7· Put it in order
v1 has the production alias and v2 has just passed testing. Put the promotion steps in the order the docs give them.
- 1.ALTER MODEL my_model VERSION v2 SET ALIAS = production;
- 2.ALTER MODEL my_model VERSION v1 UNSET ALIAS;
- 3.ALTER MODEL my_model VERSION v2 UNSET ALIAS;
First remove the alias from the old production version and clear v2's pre-release alias. Then put production on v2.
“remove the production alias from the current production version, here v1, and apply it to the new version, here v2.”Source: docs.snowflake.com
Sources3
4.Granting access to production consumers
Models are Snowflake objects, so normal role-based access control applies to them. To create a model, a role must own the schema or have CREATE MODEL on it. A model has three privileges, and the one you grant determines which kind of inference the consumer can run.
| Privilege | Allows |
|---|---|
| OWNERSHIP | Full control: managing versions, accessing artifacts, updating metadata |
| USAGE | Warehouse inference, SHOW MODELS and SHOW VERSIONS IN MODEL; no access to code, weights or artifacts |
| READ | SPCS inference, model files and metadata, SHOW MODELS and SHOW VERSIONS IN MODEL |
You can grant either privilege on a single model, on all existing models in a schema, or on future models, for example GRANT USAGE ON FUTURE MODELS IN SCHEMA <schema> TO ROLE <role>;.
Checkpoint 4 of 7· Check yourself
A role must call a model served in Snowpark Container Services but must not modify it. Which privilege fits?
USAGE covers only warehouse inference. SPCS inference needs READ, which also exposes metadata.
“The READ privilege allows grantees to use the model for SPCS inference and also see its metadata”Source: docs.snowflake.com
Sources3
5.Warehouse or Snowpark Container Services
The registry gives you one interface to two compute engines: the warehouse and Snowpark Container Services (SPCS). Which one to use depends on latency, the type of data and how you need to scale. Warehouse-deployed models have a maximum total size of 15 GB, and that limit is lower on smaller warehouse sizes. A warehouse is also a poor fit when the model requires a GPU, while SPCS suits large models and models that need GPUs.
| Pattern | Best for | Data source |
|---|---|---|
| Real-Time Inference (SPCS) | Web/mobile app backends, low-latency responses | Small inputs passed via HTTP payload |
| Native Batch Inference (SQL) | Upstream pipelines such as Dynamic Tables, dbt | Data residing in Snowflake Tables |
| Job-Based Batch (SPCS) | Images, video, audio; large historical backfills | Data residing in Snowflake Stages (Files) |
To serve a model version in SPCS, call create_service on it. It builds an image and runs the model as a service.
| Argument | Meaning |
|---|---|
| service_name | Name of the service; must be unique in the account |
| service_compute_pool | Compute pool that runs the model; must already exist (SYSTEM_COMPUTE_POOL_CPU or SYSTEM_COMPUTE_POOL_GPU may be used) |
| ingress_enabled | Must be True to call online inference from outside Snowflake |
| gpu_requests | Number of GPUs; decides CPU vs GPU for models that can run on either |
Creating a service takes time: up to 10 minutes for a CPU model and up to 20 for a GPU model, and longer if the compute pool is idle or resizing. If you send SQL inference to an SPCS-deployed model, route it through an XSMALL or SMALL warehouse. A larger warehouse sends so many concurrent requests that it can overwhelm the service.
Checkpoint 5 of 7· Check yourself
A team calls create_service on a scikit-learn model version and sets gpu_requests to "1". What happens?
Some known model types, such as scikit-learn, can only run on a CPU. Requesting GPUs for them makes the image build fail.
“the image build fails if you request GPUs”Source: docs.snowflake.com
6.Calling an externally hosted model with external functions
A model does not have to be hosted in Snowflake. An external function is a type of UDF that holds no code of its own: it calls code stored and executed outside Snowflake, known as the remote service. The remote service must accept JSON inputs, return JSON outputs and expose an HTTPS endpoint.
Snowflake does not call the remote service directly. It calls a proxy service such as Amazon API Gateway, which relays the data and the response. The security information needed to reach the proxy is stored in an API integration. External functions must be scalar, so the remote service must return exactly one row for each row it receives. Rows are sent in batches.
Checkpoint 6 of 7· Exam question
A retail data scientist's fraud model is hosted behind an Amazon SageMaker endpoint run by another team. Analysts must score rows with plain SQL inside Snowflake without moving the model. What is the MOST appropriate approach?
Correct answer: C — Create an API integration and a `CREATE EXTERNAL FUNCTION` that targets an Amazon API Gateway proxy service placed in front of the SageMaker endpoint
- A. Stored procedures do not host another team's SageMaker container, and the endpoint is not an image in a Snowflake repository. The model would simply not be reachable this way.
- B. A storage integration and external table only read files from cloud storage. They cannot invoke a hosted endpoint or return predictions.
- C. An external function calls a remote service through a cloud proxy service (such as API Gateway) using an API integration, so analysts can score rows from SQL while the model stays where it is hosted.
- D. Notification integrations send outbound messages for queues and alerts. They are not a request-response channel, so no prediction is returned to the query.
Checkpoint 7 of 7· Exam question
A team's external function `score_fraud` returns an error on every call. The remote service currently responds with the body `[0.82, 0.13, 0.55]` for a batch of three input rows. Which change to the response makes it compatible with an external function?
Correct answer: D — Return a JSON object with a `data` array that pairs each row number with its score, e.g. `{"data": [[0, 0.82], [1, 0.13], [2, 0.55]]}`
- A. Snowflake parses a JSON response; plain text lines are not a supported format, and ordering alone is not how rows are matched.
- B. Responses are not keyed by source column names. Each value must be paired with the row number it belongs to inside the `data` array.
- C. External functions are scalar: they return one value per input row, not one aggregate per batch.
- D. External functions send and expect JSON with a top-level `data` array of [row number, value] pairs, so Snowflake can match each returned value to its input row.
Sources6
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Granting USAGE on a model is enough for a role to call it in Snowpark Container Services.Why is that wrong?
USAGE covers warehouse inference only. SPCS inference requires READ, which also exposes the model's metadata.
Covered in Granting access to production consumers
2.You can pass tags to log_model when you register a model for the first time.Why is that wrong?
log_model adds a version, and the model object only exists once its first version is logged. Tags are added to the model afterwards.
Covered in Logging a model into the registry
3.create_service creates the compute pool it needs if one does not exist.Why is that wrong?
The compute pool you name must already exist. The system pools SYSTEM_COMPUTE_POOL_CPU and SYSTEM_COMPUTE_POOL_GPU can be used if the model fits them.
Covered in Warehouse or Snowpark Container Services
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“To log additional versions of the model, call log_model again with the same model_name but a different version_name.”
↩︎ Logging a model into the registry“You may specify conda_dependencies lists; the specified packages will be deployed with the model.”
↩︎ Logging a model into the registry“Models trained using Snowflake ML Functions (for example, FORECAST) do not appear in the model registry.”
↩︎ Logging a model into the registry“invoke its methods (equivalent to functions or stored procedures) to perform model operations, such as inference, in a Snowflake virtual warehouse”
↩︎ Retrieving a version and running inference“The model registry stores machine learning models as first-class schema-level objects in Snowflake.”
↩︎ Key concept“The USAGE privilege allows grantees to use the model for warehouse inference without being able to see any of its internals.”
↩︎ Exam trap 1“You cannot add tags to a model when it is added to the registry”
↩︎ Exam trap 2“The READ privilege allows grantees to use the model for SPCS inference and also see its metadata”
↩︎ Checkpoint - 2.https://docs.snowflake.com/en/developer-guide/snowflake-ml/inference/native-batch-inference-sqlOfficial docs
“Run an inference job on either a warehouse or an SPCS compute pool.”
↩︎ Retrieving a version and running inference“To run inference on SPCS instead of a warehouse, add the service_name argument to the run call:”
↩︎ Retrieving a version and running inference“To call a method of the default model, use the following syntax.”
↩︎ Retrieving a version and running inference“Memory Limits: Model size exceeds 15GB. (this limit is lower for smaller warehouse sizes).”
↩︎ Warehouse or Snowpark Container Services“Hardware Constraints: The model requires a GPU for execution.”
↩︎ Warehouse or Snowpark Container Services“Use an XSMALL or SMALL warehouse to route your inference requests to the SPCS compute pool.”
↩︎ Warehouse or Snowpark Container Services“The run method returns a dataframe that matches the type of the dataframe that you’ve specified.”
↩︎ Checkpoint - 3.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/model-managementOfficial docs
“production code only ever calls the default version of the model.”
↩︎ Promoting a version: default version and aliases“Model versions can have aliases, user-defined labels or tags that you can exclusively attach to any of a model’s versions.”
↩︎ Promoting a version: default version and aliases“Model objects have three privileges: OWNERSHIP, USAGE, and READ.”
↩︎ Granting access to production consumers“A model must always have a default version.”
↩︎ Prediction“remove the production alias from the current production version, here v1, and apply it to the new version, here v2.”
↩︎ Checkpoint - 4.https://docs.snowflake.com/en/developer-guide/snowflake-ml/inference/inference-overviewOfficial docs
“The Snowflake Model Registry provides a unified interface to both engines.”
↩︎ Warehouse or Snowpark Container Services - 5.https://docs.snowflake.com/en/developer-guide/snowflake-ml/inference/real-time-inference-rest-apiOfficial docs
“This is required to be True to call online inference from outside of Snowflake.”
↩︎ Warehouse or Snowpark Container Services“The compute pool must already exist.”
↩︎ Exam trap 3“the image build fails if you request GPUs”
↩︎ Checkpoint - 6.
“Snowflake does not call a remote service directly. Instead, Snowflake calls a proxy service, which relays the data to the remote service.”
↩︎ Calling an externally hosted model with external functions“Snowflake supports scalar external functions; the remote service must return exactly one row for each row received.”
↩︎ Calling an externally hosted model with external functions