What you will be able to do
- Generate training data from feature views with generate_training_set and know when it is materialized
- Create versioned, immutable Snowflake Datasets with generate_dataset and the Dataset API
- Read Datasets from Snowpark, TensorFlow, PyTorch, fsspec tools and SQL, with the right privileges
- Retrieve feature values for inference and explain how ML Lineage links features, datasets and models
1.Training sets from a spine DataFrame
In the Snowflake Feature Store, data for model training or scoring starts from a spine. A spine is a Snowpark DataFrame that holds entity IDs, a timestamp, label columns and any other training columns. The feature store's generate_training_set method enriches a Snowpark DataFrame that contains the source data with the derived feature values. You pass registered feature views in the features list. If you only need some features from a view, use fv.slice. For time-series features, spine_timestamp_col turns on the point-in-time lookup, so each row receives the feature values that were valid at its timestamp.
The docs describe the same pattern for every stage of the workflow: the feature table can be used to generate training datasets used for training models in Snowpark ML, or to enrich test data. The sources describe no separate validation-split API. A validation or test set is built with the same methods from its own spine and saved separately, for example as a separate Dataset or Dataset version.
By default the result is just a Snowpark DataFrame. You can pass it straight to a training function, and nothing is stored. Setting save_as to a valid table name that doesn't exist yet writes the training set to a new table. That table is not a fully governed snapshot, though: materialized tables currently don't guarantee immutability and have limited metadata support. When you need reproducible training data, the documentation points you to Snowflake Datasets instead. generate_training_set requires snowflake-ml-python 1.5.4 or later.
Checkpoint 1 of 4· Check yourself
An auditor needs the exact rows a model was trained on, and those rows must not change later. Which choice meets that requirement?
A save_as table does not guarantee immutability, and an ephemeral DataFrame is not stored at all. Dataset versions are immutable snapshots.
“Materialized tables currently don’t guarantee immutability and have limited metadata support.”Source: docs.snowflake.com
2.Snowflake Datasets: versioned, immutable snapshots
generate_dataset has a signature similar to generate_training_set, with three differences. It requires a name, it takes an optional version, and it accepts extra metadata such as desc. It also always materializes its output. The result is a Snowflake Dataset, a schema-level object built for ML. Each Dataset holds data organized into versions.
dataset: Dataset = fs.generate_dataset(
name="MY_DATASET",
spine_df=MySourceDataFrame,
features=[registered_fv],
version="v1", # optional
spine_timestamp_col="TS", # optional
spine_label_cols=["LABEL1", "LABEL2"], # optional
include_feature_view_timestamp_col=False, # optional
desc="my new dataset", # optional
)You can also build a Dataset from any Snowpark DataFrame with snowflake.ml.dataset.create_from_dataframe. Each version is an immutable, point-in-time snapshot. A Dataset object always points to one *selected* version, and reads go to that version. Creating a Dataset selects the new version automatically. create_version adds another version, and select_version returns a *new* Dataset object pointing at a different version without changing the original object. This is how you keep training, validation and retraining snapshots side by side and reproducible.
Checkpoint 2 of 4· Fill the gap
Which method adds a second version to an existing Dataset?
# Create a new version
ds2 = ds1. ? ("version2", df)
print(ds1.selected_version.name) # "version1"
print(ds2.selected_version.name) # "version2"
print(ds1.list_versions()) # ["version1", "version2"]create_version materializes a new version. select_version only switches between existing versions, and returns a new object when it does.
Source: docs.snowflake.comVersion data is stored as evenly sized Apache Parquet files. You can consume it in several ways. ds.read.to_snowpark_dataframe() returns a DataFrame for Snowpark ML Modeling. That DataFrame points to the materialized files, not to the original query. to_tf_dataset and to_torch_datapipe stream batches for TensorFlow and PyTorch. An fsspec interface supports tools such as PyArrow and Dask. From SQL, you can list files, infer the schema, or query the data directly:
LIST 'snow://dataset/<dataset_name>/versions/<dataset_version>'Datasets have their own privileges. Creating a Dataset requires the CREATE DATASET schema privilege. Adding or deleting versions requires OWNERSHIP of the Dataset. Reading requires only USAGE (or OWNERSHIP). If you set up feature store privileges with setup_feature_store or the privilege setup SQL script, Dataset privileges are set up as well. Two practical points: Datasets incur storage costs, so delete unused ones, and they don't appear in the Horizon Catalog Explorer in Snowsight.
| Aspect | generate_training_set | generate_dataset |
|---|---|---|
| Default output | Ephemeral Snowpark DataFrame | Snowflake Dataset, always materialized |
| How to persist | save_as writes a new table | name plus optional version |
| Immutability | Saved table not guaranteed immutable | Each version is an immutable snapshot |
| Metadata | Limited | Expanded, e.g. desc |
| Feeding a model | Pass the DataFrame directly | dataset.read.to_snowpark_dataframe() |
Checkpoint 3 of 4· Exam question
A fraud team registered feature view `TXN_AGG` version `1` and several models are pinned to it. The team now needs to change the aggregation window in the transformation logic. What should they do?
Correct answer: D — Register the revised DataFrame as a new version such as `2` of `TXN_AGG`, leaving version `1` intact so pinned models continue to read unchanged values.
- A. Incorrect. update_feature_view cannot change the transformation query; it only alters refresh settings and description, and silently changing version 1 would break reproducibility.
- B. Incorrect. Deleting the version drops its backing object and breaks dependent pipelines, and reusing the version string destroys the audit trail that versioning provides.
- C. Incorrect. Altering the backing object by hand bypasses Feature Store metadata and cannot change a dynamic table's query anyway; the registered definition would drift from reality.
- D. Correct. A registered feature view version is an immutable pipeline definition, so logic changes ship as a new version and consumers migrate on their own schedule.
3.Inference and end-to-end lineage
Scoring follows the same spine pattern. retrieve_feature_values takes a spine of entities to predict for, plus the registered feature views. It returns a Snowpark DataFrame with the features joined in, ready to pass to the model's predict. Because training and inference read the same registered, versioned feature views, the model sees features computed the same way in both stages. You can drop columns with exclude_columns, or keep the feature view's timestamp by setting include_feature_view_timestamp_col.
prediction_df: snowpark.DataFrame = fs.retrieve_feature_values(
spine_df=prediction_source_dataframe,
features=[registered_fv],
spine_timestamp_col="TS",
exclude_columns=[],
)These steps are also connected for governance. ML Lineage is created automatically when you use the feature store. If you produce a Dataset with generate_dataset, train on it, and log the model to the Snowflake Model Registry, the lineage graph links source data, feature views, Datasets and models. Lineage also helps at inference time: the documentation says it enables models to automatically retrieve the correct feature values at inference time, so you don't have to supply every feature input yourself.
Checkpoint 4 of 4· Put it in order
Put the Feature Store workflow steps in order, from setup to a training Dataset
- 1.Build a FeatureView from a Snowpark DataFrame that references the entity
- 2.Call generate_dataset with a spine and the registered feature view
- 3.Create or connect to the FeatureStore
- 4.Register an entity with its join keys
- 5.Register the feature view with a version
A FeatureView takes registered entities as input, must be registered before others can use it, and generate_dataset takes registered feature views.
“Once a feature view has been completely defined, you can register it in the feature store using the feature store’s register_feature_view method”Source: docs.snowflake.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Setting save_as on generate_training_set gives you an immutable, reproducible training snapshot.Why is that wrong?
save_as writes an ordinary table that is not guaranteed immutable and has limited metadata. Use generate_dataset when you need immutability.
Covered in Training sets from a spine DataFrame
2.Reading a Dataset only needs access to the schema, and adding versions only needs USAGE.Why is that wrong?
Reading needs USAGE or OWNERSHIP on the Dataset itself, and adding or deleting versions needs OWNERSHIP.
Covered in Snowflake Datasets: versioned, immutable snapshots
Practise it for real
Go from a registered feature view to a versioned training Dataset and a scoring DataFrame
1.Connect to your feature store with FeatureStore(session=..., database=..., name=..., default_warehouse=...) using the default creation mode.
Why: The default mode connects to an existing store and fails if the store doesn't exist.
You should see: An fs object, or an exception if the feature store schema does not exist.
2.Run fs.list_feature_views().show() and fetch one view with fs.get_feature_view(name=..., version=...).
Why: Registered, versioned feature views are the inputs that the dataset methods accept.
You should see: A Snowpark DataFrame listing the registered feature views, plus a FeatureView object.
3.Call fs.generate_dataset(name="MY_DATASET", spine_df=..., features=[registered_fv], version="v1", spine_timestamp_col="TS").
Why: generate_dataset always materializes an immutable, versioned snapshot.
You should see: A Dataset whose selected version is v1.
4.Run LIST 'snow://dataset/MY_DATASET/versions/v1' in SQL.
Why: This shows how the version is stored on disk.
You should see: Parquet files belonging to the dataset version.
5.Call dataset.read.to_snowpark_dataframe(), then fs.retrieve_feature_values(spine_df=..., features=[registered_fv], spine_timestamp_col="TS") for a new spine.
Why: The first call gives you training input. The second gives you scoring input from the same feature views.
You should see: Two Snowpark DataFrames, one read from the Dataset files and one with features joined to the new spine.
Stuck? Get a nudge
If step 4 is denied, check that your role has USAGE or OWNERSHIP on the Dataset.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“which enriches a Snowpark DataFrame that contains the source data with the derived feature values.”
↩︎ Training sets from a spine DataFrame“Requested features are retrieved for the list of entity IDs, with point-in-time correctness with respect to the provided time stamp.”
↩︎ Training sets from a spine DataFrame“generate_dataset always materializes the result.”
↩︎ Snowflake Datasets: versioned, immutable snapshots“you can retrieve the feature view from the feature store and pass it to your model for prediction”
↩︎ Inference and end-to-end lineage“Materialized tables currently don’t guarantee immutability and have limited metadata support.”
↩︎ Exam trap 1“Training sets are ephemeral by default; they exist only as Snowpark DataFrames and are not materialized.”
↩︎ Prediction“Materialized tables currently don’t guarantee immutability and have limited metadata support.”
↩︎ Checkpoint - 2.
“The feature table can be used to generate training datasets used for training models in Snowpark ML, or to enrich test data”
↩︎ Training sets from a spine DataFrame“ML Lineage is automatically created when the feature store is used.”
↩︎ Inference and end-to-end lineage - 3.
“Each version is an immutable, point-in-time snapshot of the data managed by the Dataset.”
↩︎ Snowflake Datasets: versioned, immutable snapshots“Dataset version data is stored as evenly sized files in the Apache Parquet format.”
↩︎ Snowflake Datasets: versioned, immutable snapshots“Creating Datasets requires the CREATE DATASET schema-level privilege.”
↩︎ Snowflake Datasets: versioned, immutable snapshots“Datasets incur storage costs. Delete unused datasets to minimize costs.”
↩︎ Snowflake Datasets: versioned, immutable snapshots“automatically completing the ML Lineage graph linking source data, feature views, datasets, and models for full end-to-end governance.”
↩︎ Inference and end-to-end lineage“Reading from a Dataset requires only the USAGE privilege on the Dataset (or OWNERSHIP).”
↩︎ Exam trap 2
Also cited
“Once a feature view has been completely defined, you can register it in the feature store using the feature store’s register_feature_view method”
↩︎ Checkpoint