What you will be able to do
- Log a model version with log_model and declare its metadata, metrics and dependencies
- Package pickled files, in-memory models and helper code into a CustomModel with ModelContext and code_paths
- Export a logged version's files and know which registry limits a logging job can hit
- Validate a model artifact before and during registration using local tests, signatures and target_platforms
- Check that a target environment matches the one a model was logged in before loading or deploying it
- Roll back to a previous model version when performance degrades, without editing any version
Key concept
Model version as an immutable registered artifact — Each call to log_model serializes a Python model and stores it as one named version of a schema-level model object. After that, the code and weights are fixed. Only metadata such as comments, metrics, aliases and tags can change, so promotion and rollback work by pointing at a different version, never by editing one.
1.Logging a model version with metadata and dependencies
You can use any Snowflake schema as a registry, and it needs no setup. Snowflake recommends a dedicated schema such as ML.REGISTRY. In Python you open it with Registry(session=session, database_name="ML", schema_name="REGISTRY"), then add models by calling log_model. That call serializes the Python model object and creates a Snowflake model object from it. The first call for a given model_name creates the model. Each later call with the same name and a new version_name adds a version. The name and version together must be unique in the schema. If you leave out version_name, Snowflake generates a human-readable one.
from snowflake.ml.model import task, type_hints
mv = reg.log_model(clf,
model_name="my_model",
version_name="v1",
conda_dependencies=["scikit-learn"],
comment="My awesome ML model",
metrics={"score": 96},
sample_input_data=train_features,
task=task.Task.TABULAR_BINARY_CLASSIFICATION)You can add metadata after logging too. mv.set_metric("accuracy", 0.769) in Python, or ALTER MODEL … MODIFY VERSION … SET METADATA in SQL, stores metrics on the version. Those values then appear in INFORMATION_SCHEMA.MODEL_VERSIONS, where you can rank every version by accuracy. Tags are different: they belong to the model, not the version, so you can only set them after the first version exists. Aliases are another kind of metadata. An alias is a user-defined label, such as production, that you attach to one version with ALTER MODEL … VERSION … SET ALIAS. It can be used wherever a version name is accepted, and it can be moved to a different version later because it is not part of the version's code. Page 2 covers aliases in detail. You declare dependencies with the arguments below.
| Argument | What it controls |
|---|---|
| conda_dependencies | Conda packages in "[channel::]package [operator version]" format. With no channel, the Snowflake channel is assumed on a warehouse and conda-forge on SPCS |
| pip_requirements | PyPI package specs. Models that target a warehouse also need artifact_repository_map |
| artifact_repository_map | Maps "pip" to a repository, e.g. snowflake.snowpark.pypi_shared_repository |
| python_version | Python version the model runs under. None means the latest version available in the warehouse |
| code_paths | Directories of code to import when loading or deploying the model |
Checkpoint 1 of 8· Check yourself
A model is logged with conda_dependencies=["xgboost"] and no channel, then run in a virtual warehouse. Which channel supplies the package?
When no channel is given, the Snowflake channel is assumed for warehouse execution. conda-forge is the assumption only for SPCS.
“If you do not specify a channel, the Snowflake channel is assumed when the model runs on a warehouse.”Source: docs.snowflake.com
Registration has hard limits. The registry's limits table says a model can hold at most 1000 versions, yet the same page also says a model can have unlimited versions. The sources do not reconcile these two statements, so for planning treat 1000 as the documented limit and check the current docs before relying on either number. Each version allows at most 10 methods, 500 arguments per method, 100 KB of metadata including metrics, and 15 GB of total size for warehouse-deployed models. Its config files, including the conda.yml that log_model generates, can total at most 250 KB. A nightly retraining job that logs a new version every run could reach the documented version cap unless it also removes old versions.
Checkpoint 2 of 8· Exam question
A churn model is served from `ML.PROD.CHURN`, and every scoring query uses `WITH m AS MODEL ML.PROD.CHURN VERSION PROD SELECT m!PREDICT(...)`. Version V7, currently carrying the `PROD` alias, shows degraded precision, while V6 is the last version that passed validation. Which rollback leaves all client SQL untouched?
Correct answer: A — Remove `PROD` from V7 and assign it to V6 with `ALTER MODEL ... VERSION ... SET ALIAS`, so existing queries resolve to V6 immediately.
- A. Alias reassignment is the intended rollback path. Because clients reference the alias rather than a version name, repointing it restores V6 without touching any client code or artifacts.
- B. Dropping V7 destroys the evidence needed to diagnose the regression and removes a rollback target. Re-logging V6 under a misleading name also corrupts version lineage, and the alias would still need handling.
- C. Hard-coding V6 into each query couples clients to a version name again, which is the problem aliases solve. It needs coordinated deployments and makes the next rollback just as expensive.
- D. A new model object has a different fully qualified name, so every client query would still need editing. It also splits history and privileges across two model objects rather than reusing the version stored in the original.
2.Custom models, pickle files and extracting artifacts
If a model type has no built-in support, you wrap it in a subclass of custom_model.CustomModel. Its artifacts go in a custom_model.ModelContext, which takes keyword arguments. Each value is either an in-memory object of a supported model type or a string path to a file, such as a pickle, a JSON config or a parameter file. Files and serialized models are packaged with the version. If you pass a supported model object, such as an XGBoost model, Snowflake serializes it for you. Files you load yourself in __init__.
model_context = custom_model.ModelContext(
my_file_path='/path/to/file.pkl',
)
# Define a custom model class that loads the pickled object
class ExampleBringYourOwnModel(custom_model.CustomModel):
def __init__(self, context: custom_model.ModelContext) -> None:
super().__init__(context)
# Use 'my_file_path' key from the context to load the pickled object
with open(self.context['my_file_path'], 'rb') as f:
self.obj = pickle.load(f)Inside the class, always read models through self.context[...]. If you assign an in-memory model straight to an attribute, the closure captures a second copy and the serialized artifact gets much larger. You add helper modules with code_paths. A plain string path copies a file or directory, and its last component becomes the import name. A CodePath(root, filter=...) object includes only part of a directory tree, for example so tests are left out. A separate argument, ext_modules, lists external modules to pickle with the model.
Checkpoint 3 of 8· Check yourself
A custom model's __init__ sets self.model = xgb_model, using an in-memory object instead of reading it from the context. What is the documented consequence?
Reading the object directly captures a second copy in a closure, so serialization produces much larger files. Read it through self.context instead.
“Accessing the model directly captures a second copy of the model in a closure, which results in significantly larger model files during serialization.”Source: docs.snowflake.com
Checkpoint 4 of 8· Fill the gap
Which log_model argument packages the helper module so the model can import it?
mv = reg.log_model(
my_model,
model_name="my_model",
version_name="v1",
? =["src/mymodule"], # import with: import mymodule
)code_paths copies the listed files or directories into the model, and the last path component becomes the importable module name.
Source: docs.snowflake.comArtifacts can also come back out of the registry. mv.export("~/mymodel/") writes a version's files to a local directory and creates the directory if it is missing. Access to artifacts depends on privilege, and the docs word it in two ways. The model-management page says OWNERSHIP includes accessing artifacts, and that READ allows model files and metadata. The registry overview page says you can view or download the artifacts of models you own, and that ownership is required. The Snowsight Files tab, which lists a version's underlying artifacts, is available with either OWNERSHIP or READ. OWNERSHIP is the one privilege every page agrees gives artifact access. A role with only USAGE can run warehouse inference but cannot see code, weights or other artifacts.
Checkpoint 5 of 8· Exam question
A security team needs to review the pickle files bundled in version V3 of a custom model. The analyst role holds `USAGE` on the model, but `LIST 'snow://model/churn/versions/V3/'` is rejected. What does the role need?
Correct answer: B — `OWNERSHIP` on the model, because access to stored artifacts through `snow://` URLs is restricted to the role that owns the model.
- A. `READ` allows SPCS inference and viewing metadata, but the registry documentation reserves access to the stored artifact files for the model owner, so this grant alone still fails.
- B. Artifact access requires ownership of the model. `USAGE` only supports warehouse inference and SHOW commands, so the owner role (or a role granted ownership) is what makes the listing succeed.
- C. Future grants only apply to models created later and still confer `USAGE`, which exposes no internals. It would not change the privilege level this role holds on the existing model.
- D. `CREATE MODEL` lets a role register new models in the schema. It grants no rights on existing model objects, and listing artifacts is not treated as a creation operation.
3.Validating artifact integrity before and during registration
The registry has several integrity gates. First, the model object passed to log_model must be serializable, which the docs call "pickleable". Second, every model except Snowpark ML models, MLFlow models and Hugging Face pipelines needs either sample_input_data or explicit signatures. These define the feature names and types used to validate input. Third, target_platforms checks whether the model can run where you say it will. If WAREHOUSE is listed and dependencies, GPU needs or size make the model unrunnable there, log_model() fails instead of registering a broken version. Finally, a version cannot be modified after logging, so the artifact you validated is the artifact that gets served.
You can also check artifacts after registration, for example before copying them to another environment. LIST 'snow://model/my_model/versions/V3/'; returns each artifact with its size and an md5 column. You can record those values at registration time and compare them with a later listing or with the files you retrieve with GET. The docs show the md5 column but do not describe a built-in verification procedure, so the comparison is a practice you apply yourself. The docs also note that artifact names and organization vary by model type and might change. Treat code that a model brings with it as a trust decision. When you import a model through Snowsight, the Trust remote code option lets the model download arbitrary code, which the docs call a security risk. Only allow it for models you have thoroughly evaluated and trust.
The PyCaret walkthrough shows the full sequence for a custom model. Test it locally, infer a signature from the test input and output, log it with pinned dependencies and relax_version turned off, confirm it is there with show_models(), then call run. The custom methods must also handle multi-row DataFrames. Server-side batching can combine single-record requests into one frame.
custom_mv = snowml_registry.log_model(
my_pycaret_model,
model_name="my_pycaret_best_model",
version_name="version_1",
conda_dependencies=["pycaret==3.0.2", "scipy==1.11.4", "joblib==1.2.0"],
options={"relax_version": False},
signatures={"predict": predict_signature},
comment = 'My PyCaret classification experiment using the CustomModel API'
)Checkpoint 6 of 8· Put it in order
Put the documented PyCaret custom-model workflow in order
- 1.Infer a model signature from the test input and output
- 2.Run the custom model's predict locally on test data
- 3.Call show_models() to verify the model is in the registry
- 4.Log the model with log_model, passing the signature
- 5.Call run on the version to get predictions
The walkthrough tests locally, builds the signature from that test, logs, verifies with show_models, then runs inference.
“To verify that the model is available in the Model Registry, use show_models function.”Source: docs.snowflake.com
Checkpoint 7 of 8· Check yourself
A GPU-only custom model is logged with target_platforms=["WAREHOUSE"]. What happens?
The WAREHOUSE target is checked when the model is logged. A model that cannot run there makes log_model() fail.
“and the model is not runnable in the warehouse (due to dependencies, gpu requirement, model size etc), log_model() fails.”Source: docs.snowflake.com
Compatibility across environments is the other half of validation. A model loaded with mv.load() works properly only when the target Python environment, meaning the interpreter version and every library version, is identical to the one it was logged from. Passing force=True to the load call loads it even when the environment differs, so use it knowingly, because it skips that protection. To make a dev, test or prod environment match the one hosting the model, download the version's conda.yml from the registry and build a conda environment from it. At logging time, python_version and resource_constraint (for example an architecture such as x86) fix the runtime, and target_platforms fails the log if the model cannot run in a warehouse. Check these before promoting a version to another environment.
session.file.get("snow://model/<modelName>/versions/<versionName>/runtimes/python_runtime/env/conda.yml", "~/")4.Rolling back a version when performance degrades
Because a version never changes after logging, rolling back means pointing production at an older version, not editing the new one. Keep previous production versions so there is something to return to. How you roll back depends on how you promoted. With the default-version scheme, set the default back to the earlier version, and callers of my_model!predict(...) pick it up with no code change. With the alias scheme, unset the alias on the bad version and set it on the earlier one, since an alias can sit on only one version at a time. Callers that use MODEL(my_model, production) follow the alias. Do not drop the bad version first. The default version cannot be dropped, so change the default before you drop anything. Set a policy for how many old versions to keep and for how long.
ALTER MODEL prod_db.prod_schema.prod_model SET DEFAULT_VERSION = V1;Checkpoint 8 of 8· Check yourself
Version V2 is the default and its accuracy has dropped in production. V1 is still in the model. What is the documented way to roll back?
Versions are immutable and the default version cannot be dropped, so you roll back by setting the default to the previous version.
“which you can do by setting the default version to a previous version, as shown below.”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Pinning ==x.y.z in conda_dependencies guarantees that exact version is deployed.Why is that wrong?
relax_version defaults to True and loosens the pins. Pass options={"relax_version": False} to keep exact versions.
Covered in Validating artifact integrity before and during registration
2.In a CustomModel, storing the in-memory model on self is equivalent to reading it from the context.Why is that wrong?
Direct assignment captures a second copy in a closure and inflates the serialized artifact. Always use self.context[key].
Covered in Custom models, pickle files and extracting artifacts
3.To roll back a bad default version, drop it first and the previous version takes over.Why is that wrong?
The default version cannot be dropped. Set the default to the earlier version first, then drop the unneeded one if you want.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“To log additional versions of the model, call log_model again with the same model_name but a different version_name.”
↩︎ Logging a model version with metadata and dependencies“Maximum of 10 methods Maximum of 500 arguments per method”
↩︎ Logging a model version with metadata and dependencies“Models | - Maximum of 1000 versions”
↩︎ Logging a model version with metadata and dependencies“A model can have unlimited versions, each identified by a string.”
↩︎ Logging a model version with metadata and dependencies“Use mv.export to export a model’s files to a local directory; the directory is created if it does not exist”
↩︎ Custom models, pickle files and extracting artifacts“Artifacts cannot be modified, but you can view or download the artifacts of models you own.”
↩︎ Custom models, pickle files and extracting artifacts“Must be serializable (“pickleable”).”
↩︎ Validating artifact integrity before and during registration“LIST 'snow://model/my_model/versions/V3/';”
↩︎ Validating artifact integrity before and during registration“Specify force=True in the load call to force the model to be loaded even if the environment is different.”
↩︎ Validating artifact integrity before and during registration“This can be used to ensure the model runs in a warehouse with the necessary architecture.”
↩︎ Validating artifact integrity before and during registration“After registration, the model itself cannot be modified (although you can change its metadata).”
↩︎ Key concept“relax_version: whether to relax the version constraints of the dependencies. This replaces version specifiers like ==x.y.z with specifiers like <=x.y, <(x+1). Default: True.”
↩︎ Exam trap 1“If you do not specify a channel, the Snowflake channel is assumed when the model runs on a warehouse.”
↩︎ Checkpoint“relax_version: whether to relax the version constraints of the dependencies. This replaces version specifiers like ==x.y.z with specifiers like <=x.y, <(x+1). Default: True.”
↩︎ Prediction“and the model is not runnable in the warehouse (due to dependencies, gpu requirement, model size etc), log_model() fails.”
↩︎ Checkpoint - 2.
“An alias is an alternative name that can be easily reassigned.”
↩︎ Logging a model version with metadata and dependencies“A version can have at most one alias.”
↩︎ Rolling back a version when performance degrades - 3.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/model-managementOfficial docs
“The view for models and their versions is INFORMATION_SCHEMA.MODEL_VERSIONS.”
↩︎ Logging a model version with metadata and dependencies“Roles with only USAGE cannot access the model code, weights, or other artifacts.”
↩︎ Custom models, pickle files and extracting artifacts“allowing SPCS inference (prediction), model files, metadata and use of the SHOW MODELS”
↩︎ Custom models, pickle files and extracting artifacts“which you can do by setting the default version to a previous version, as shown below.”
↩︎ Rolling back a version when performance degrades - 4.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/bring-your-own-model-typesOfficial docs
“Set the supported model object directly in the context (e.g., base_model = my_xgb_model) and it is serialized automatically.”
↩︎ Custom models, pickle files and extracting artifacts“use the sample data to infer a model signature for input validation”
↩︎ Validating artifact integrity before and during registration“Methods decorated with @custom_model.inference_api should always be written to work on multi-row dataframe.”
↩︎ Validating artifact integrity before and during registration“In your custom model class, always access model objects through the model context.”
↩︎ Exam trap 2“Accessing the model directly captures a second copy of the model in a closure, which results in significantly larger model files during serialization.”
↩︎ Checkpoint“To verify that the model is available in the Model Registry, use show_models function.”
↩︎ Checkpoint - 5.
“This page is only available if the user has either OWNERSHIP or READ privilege on the model.”
↩︎ Custom models, pickle files and extracting artifacts“Allowing models to download arbitrary code should be considered a security risk.”
↩︎ Validating artifact integrity before and during registration - 6.
“You cannot drop the default version of a model.”
↩︎ Rolling back a version when performance degrades“You cannot drop the default version of a model.”
↩︎ Exam trap 3