What you will be able to do
- Log a model and its later versions with Registry.log_model, knowing what makes a version unique
- Tell apart comments, tags and metrics: what each is for and what it attaches to
- Query version metadata through INFORMATION_SCHEMA.MODEL_VERSIONS and connect a version to the immutable data it was trained on
Key concept
Model version — In the Snowflake Model Registry, a model is a schema-level object that holds named versions. Each log_model call adds one version. Training, evaluation, promotion and rollback all happen at the version level.
1.Logging a model creates a version
The Model Registry stores models as first-class, schema-level Snowflake objects. You don't need a special registry database. Any Snowflake schema can serve as a registry, although Snowflake recommends creating a dedicated schema such as ML.REGISTRY. In Python you open the registry with Registry(session=..., database_name=..., schema_name=...). It returns a handle for logging new models and getting references to existing ones. The main classes are Registry (manages models within a schema), Model (represents a model) and ModelVersion (represents one version of a model).
Adding a model to the registry is called *logging* it. log_model serializes the Python model object and creates the Snowflake model object, together with any metadata you pass in the same call. A model object only comes into existence when its first version is logged. To retrain and register a new version, you call log_model again with the same model_name and a new version_name. The pair must be unique in the schema, and the model name can't be changed after logging. The call below logs version v1 with a pip-free dependency list (conda_dependencies), a comment, an initial metric and sample input data.
from snowflake.ml.model import task, type_hints
mv = reg.log_model(clf,
model_name="my_model",
version_name="v1",
conda_dependencies=["scikit-learn"],
comment="My awesome ML model",
metrics={"score": 96},
sample_input_data=train_features,
task=task.Task.TABULAR_BINARY_CLASSIFICATION)The registry has limits worth knowing. A model can hold at most 1,000 versions. Each version can have at most 10 methods and at most 100 KB of metadata, metrics included. For warehouse-deployed models, the total model size is capped at 15 GB. Built-in type support covers scikit-learn, XGBoost, LightGBM, Prophet, CatBoost, PyTorch, TensorFlow, Keras, Hugging Face pipelines, Sentence Transformer and MLflow pyfunc models. Models trained with Snowflake ML Functions such as FORECAST do not appear in the registry.
Checkpoint 1 of 5· Check yourself
A nightly job retrains a churn classifier and needs to register the result alongside the model already in the registry as CHURN / v1. What should the job do?
New versions go under the same model name with a different version name. Reusing v1 fails because the name–version pair must be unique in the schema. A separate model name would split the version history, so default-version promotion would no longer work across them.
“To log additional versions of the model, call log_model again with the same model_name but a different version_name.”Source: docs.snowflake.com
Sources1
2.Metadata: comments, tags and metrics
Once a model exists, the registry gives you three kinds of metadata, and they work differently. A comment (also available as description) is free text. Both models and model versions can have one, and you can set it from Python or with ALTER MODEL. Tags are standard Snowflake object tags applied to the model. They record purpose, algorithm, training data set, lifecycle stage or anything else you choose. Metrics are key-value pairs on a specific *version*. Their values can be any JSON-serializable Python object.
The big practical difference is up-front definition. A tag name has to exist as a schema-level object, created with CREATE TAG, before you can apply it. In Python you then use set_tag, get_tag, show_tags and unset_tag on the model. Metric names need no setup. You set them at log time through metrics= or later with set_metric, read them back with show_metrics, and remove them with delete_metric.
# scalar metric
mv.set_metric("test_accuracy", test_accuracy)
# hierarchical (dictionary) metric
mv.set_metric("evaluation_info", {"dataset_used": "my_dataset", "accuracy": test_accuracy, "f1_score": f1_score})
# multivalent (matrix) metric
mv.set_metric("confusion_matrix", test_confusion_matrix)Because tags are ordinary Snowflake object tags, the general tagging controls apply to them. ALLOWED_VALUES limits a tag to a fixed list of strings, which suits a lifecycle-stage tag where only a few values make sense. Each tag value can be up to 256 characters long.
CREATE TAG cost_center ALLOWED_VALUES 'finance', 'engineering';| Metadata | Attached to | Must be defined first? | Python API |
|---|---|---|---|
| Comment | Model or model version | No | comment / description attribute |
| Tag | Model | Yes, with CREATE TAG | set_tag, get_tag, show_tags, unset_tag |
| Metric | Model version | No | set_metric, show_metrics, delete_metric |
Checkpoint 2 of 5· Match them up
Match each piece of registry metadata to its defining behaviour
Tap a term, then the definition that fits it.
Tags are governed Snowflake objects, so they have to be created first. Metrics are lightweight version-level key-value pairs that need no setup. Comment and description are synonyms.
“Unlike tags, metric names and possible values do not need to be defined in advance.”Source: docs.snowflake.com
3.Querying metadata and tracing versions to their data
Models are Snowflake objects, so they appear in the Information Schema. A single view, INFORMATION_SCHEMA.MODEL_VERSIONS, covers both models and versions. Version information is a superset of model information, so there is no separate MODEL view. Metrics are stored as version metadata, which means you can also write them in SQL with ALTER MODEL ... MODIFY VERSION ... SET METADATA.
Checkpoint 3 of 5· Fill the gap
Which keyword completes this statement, which records an accuracy metric on version v1 in SQL?
ALTER MODEL my_model MODIFY VERSION v1
SET ? = '{"metric": {"accuracy": 0.769}}';Metrics live in the version's metadata, so SQL sets them through SET METADATA on MODIFY VERSION. Tags and aliases are separate mechanisms.
Source: docs.snowflake.comOnce every version carries the same metric, you can rank all versions of all models in a schema with a single query. That makes it a convenient way to compare a retrained candidate with the versions already registered.
SELECT
catalog_name,
schema_name,
model_name,
model_version_name,
metadata:metric:accuracy AS accuracy,
comment,
owner,
functions,
created_on,
last_altered_on
FROM my_database.INFORMATION_SCHEMA.MODEL_VERSIONS
ORDER BY accuracy DESC;A metric records how well a version performed. It doesn't preserve the data the version was trained on. Snowflake Datasets do that. They are schema-level objects organized into versions, and each version is a materialized, immutable snapshot of the data. One of their documented uses is tracking the lineage used to create an ML model. If you train from a Dataset version and record its name on the model version, for example in a metric such as evaluation_info with a dataset_used key, the training data stays identifiable even after the source table keeps changing.
Checkpoint 4 of 5· Check yourself
An auditor asks which exact rows trained version v2, but the source table has been updated daily since training. Which practice would have made the question answerable?
Only an immutable snapshot keeps the exact training data. A table name, date or timestamp points to a table whose contents have changed since training.
“Each version holds a materialized snapshot of your data with guaranteed immutability”Source: docs.snowflake.com
Checkpoint 5 of 5· Exam question
A data scientist validated churn model version V3 against holdout data. Dozens of existing SQL jobs call `CHURN_MODEL!PREDICT(...)` without naming a version. Which action moves all of them to V3 with no SQL edits?
Correct answer: C — Set the model's default version to V3 with `ALTER MODEL CHURN_MODEL SET DEFAULT_VERSION = V3`, which unqualified calls resolve to.
- A. Incorrect: dropping the model destroys version history, comments and metrics needed for rollback, and it is unnecessary to change the default.
- B. Incorrect: version names are unique within a model and logged artifacts are immutable, so logging V2 again fails rather than replacing weights.
- C. Correct: unqualified `MODEL!PREDICT` calls resolve to the default version, so changing the default switches every caller at once.
- D. Incorrect: unqualified calls bind to the default version rather than the newest, and editing each job is exactly the manual work to avoid.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Tags are set per model version, like metrics, so each version can carry its own tag values.Why is that wrong?
Tags are attributes of the model object, while metrics belong to individual versions. A tag that stores a version name, such as live_version, is set once on the model and updated as versions change.
Covered in Metadata: comments, tags and metrics
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The combination of model name and version must be unique in the schema.”
↩︎ Logging a model creates a version“Thus, any Snowflake schema can be used as a registry.”
↩︎ Logging a model creates a version“You must define the names of all tags (and potentially their possible values) first by using CREATE TAG.”
↩︎ Metadata: comments, tags and metrics“Metrics are key-value pairs used to track prediction accuracy and other model version characteristics.”
↩︎ Metadata: comments, tags and metrics“Logging a model actually logs a specific version of the model.”
↩︎ Key concept“because tags are attributes of the model, and log_model adds a specific model version”
↩︎ Exam trap 1“To log additional versions of the model, call log_model again with the same model_name but a different version_name.”
↩︎ Checkpoint“Unlike tags, metric names and possible values do not need to be defined in advance.”
↩︎ Checkpoint - 2.
“Users cannot assign a value to a tag unless the value is in the defined list.”
↩︎ Metadata: comments, tags and metrics - 3.https://docs.snowflake.com/en/developer-guide/snowflake-ml/model-registry/model-managementOfficial docs
“Model version information is a superset of the information for models, so there is no separate MODEL view.”
↩︎ Querying metadata and tracing versions to their data - 4.
“You need to track the lineage used to create an ML model.”
↩︎ Querying metadata and tracing versions to their data“Each version holds a materialized snapshot of your data with guaranteed immutability”
↩︎ Checkpoint