CertSafari
    Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01)· Lessons

    Domain 1 · Lesson 2/17

    Feature Store Datasets for Training, Validation and Inference

    Implement Snowflake Feature Store architecture and management.

    9 min read
    4% of exam
    4 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Generate training data from feature views with generate_training_set and know when it is materialized
    • Create versioned, immutable Snowflake Datasets with generate_dataset and the Dataset API
    • Read Datasets from Snowpark, TensorFlow, PyTorch, fsspec tools and SQL, with the right privileges
    • Retrieve feature values for inference and explain how ML Lineage links features, datasets and models

    1.Training sets from a spine DataFrame

    In the Snowflake Feature Store, data for model training or scoring starts from a spine. A spine is a Snowpark DataFrame that holds entity IDs, a timestamp, label columns and any other training columns. The feature store's generate_training_set method enriches a Snowpark DataFrame that contains the source data with the derived feature values. You pass registered feature views in the features list. If you only need some features from a view, use fv.slice. For time-series features, spine_timestamp_col turns on the point-in-time lookup, so each row receives the feature values that were valid at its timestamp.

    The docs describe the same pattern for every stage of the workflow: the feature table can be used to generate training datasets used for training models in Snowpark ML, or to enrich test data. The sources describe no separate validation-split API. A validation or test set is built with the same methods from its own spine and saved separately, for example as a separate Dataset or Dataset version.

    By default the result is just a Snowpark DataFrame. You can pass it straight to a training function, and nothing is stored. Setting save_as to a valid table name that doesn't exist yet writes the training set to a new table. That table is not a fully governed snapshot, though: materialized tables currently don't guarantee immutability and have limited metadata support. When you need reproducible training data, the documentation points you to Snowflake Datasets instead. generate_training_set requires snowflake-ml-python 1.5.4 or later.

    Checkpoint 1 of 4· Check yourself

    An auditor needs the exact rows a model was trained on, and those rows must not change later. Which choice meets that requirement?

    Sources12

    2.Snowflake Datasets: versioned, immutable snapshots

    generate_dataset has a signature similar to generate_training_set, with three differences. It requires a name, it takes an optional version, and it accepts extra metadata such as desc. It also always materializes its output. The result is a Snowflake Dataset, a schema-level object built for ML. Each Dataset holds data organized into versions.

    Generating a versioned Dataset from a registered feature viewpython
    dataset: Dataset = fs.generate_dataset(
        name="MY_DATASET",
        spine_df=MySourceDataFrame,
        features=[registered_fv],
        version="v1",                               # optional
        spine_timestamp_col="TS",                   # optional
        spine_label_cols=["LABEL1", "LABEL2"],      # optional
        include_feature_view_timestamp_col=False,   # optional
        desc="my new dataset",                      # optional
    )

    You can also build a Dataset from any Snowpark DataFrame with snowflake.ml.dataset.create_from_dataframe. Each version is an immutable, point-in-time snapshot. A Dataset object always points to one *selected* version, and reads go to that version. Creating a Dataset selects the new version automatically. create_version adds another version, and select_version returns a *new* Dataset object pointing at a different version without changing the original object. This is how you keep training, validation and retraining snapshots side by side and reproducible.

    Checkpoint 2 of 4· Fill the gap

    Which method adds a second version to an existing Dataset?

    # Create a new version
    ds2 = ds1. ? ("version2", df)
    print(ds1.selected_version.name)  # "version1"
    print(ds2.selected_version.name)  # "version2"
    print(ds1.list_versions())        # ["version1", "version2"]

    Version data is stored as evenly sized Apache Parquet files. You can consume it in several ways. ds.read.to_snowpark_dataframe() returns a DataFrame for Snowpark ML Modeling. That DataFrame points to the materialized files, not to the original query. to_tf_dataset and to_torch_datapipe stream batches for TensorFlow and PyTorch. An fsspec interface supports tools such as PyArrow and Dask. From SQL, you can list files, infer the schema, or query the data directly:

    Listing the files in a Dataset version from SQLsql
    LIST 'snow://dataset/<dataset_name>/versions/<dataset_version>'

    Datasets have their own privileges. Creating a Dataset requires the CREATE DATASET schema privilege. Adding or deleting versions requires OWNERSHIP of the Dataset. Reading requires only USAGE (or OWNERSHIP). If you set up feature store privileges with setup_feature_store or the privilege setup SQL script, Dataset privileges are set up as well. Two practical points: Datasets incur storage costs, so delete unused ones, and they don't appear in the Horizon Catalog Explorer in Snowsight.

    generate_training_set vs generate_dataset
    Aspectgenerate_training_setgenerate_dataset
    Default outputEphemeral Snowpark DataFrameSnowflake Dataset, always materialized
    How to persistsave_as writes a new tablename plus optional version
    ImmutabilitySaved table not guaranteed immutableEach version is an immutable snapshot
    MetadataLimitedExpanded, e.g. desc
    Feeding a modelPass the DataFrame directlydataset.read.to_snowpark_dataframe()

    Checkpoint 3 of 4· Exam question

    A fraud team registered feature view `TXN_AGG` version `1` and several models are pinned to it. The team now needs to change the aggregation window in the transformation logic. What should they do?

    Sources13

    3.Inference and end-to-end lineage

    Scoring follows the same spine pattern. retrieve_feature_values takes a spine of entities to predict for, plus the registered feature views. It returns a Snowpark DataFrame with the features joined in, ready to pass to the model's predict. Because training and inference read the same registered, versioned feature views, the model sees features computed the same way in both stages. You can drop columns with exclude_columns, or keep the feature view's timestamp by setting include_feature_view_timestamp_col.

    Retrieving feature values for predictionpython
    prediction_df: snowpark.DataFrame = fs.retrieve_feature_values(
        spine_df=prediction_source_dataframe,
        features=[registered_fv],
        spine_timestamp_col="TS",
        exclude_columns=[],
    )

    These steps are also connected for governance. ML Lineage is created automatically when you use the feature store. If you produce a Dataset with generate_dataset, train on it, and log the model to the Snowflake Model Registry, the lineage graph links source data, feature views, Datasets and models. Lineage also helps at inference time: the documentation says it enables models to automatically retrieve the correct feature values at inference time, so you don't have to supply every feature input yourself.

    Checkpoint 4 of 4· Put it in order

    Put the Feature Store workflow steps in order, from setup to a training Dataset

    1. 1.Build a FeatureView from a Snowpark DataFrame that references the entity
    2. 2.Call generate_dataset with a spine and the registered feature view
    3. 3.Create or connect to the FeatureStore
    4. 4.Register an entity with its join keys
    5. 5.Register the feature view with a version

    Sources132

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Setting save_as on generate_training_set gives you an immutable, reproducible training snapshot.Why is that wrong?

      save_as writes an ordinary table that is not guaranteed immutable and has limited metadata. Use generate_dataset when you need immutability.

      Covered in Training sets from a spine DataFrame

    2. 2.Reading a Dataset only needs access to the schema, and adding versions only needs USAGE.Why is that wrong?

      Reading needs USAGE or OWNERSHIP on the Dataset itself, and adding or deleting versions needs OWNERSHIP.

      Covered in Snowflake Datasets: versioned, immutable snapshots

    Practise it for real

    Go from a registered feature view to a versioned training Dataset and a scoring DataFrame

    1. 1.Connect to your feature store with FeatureStore(session=..., database=..., name=..., default_warehouse=...) using the default creation mode.

      Why: The default mode connects to an existing store and fails if the store doesn't exist.

      You should see: An fs object, or an exception if the feature store schema does not exist.

    2. 2.Run fs.list_feature_views().show() and fetch one view with fs.get_feature_view(name=..., version=...).

      Why: Registered, versioned feature views are the inputs that the dataset methods accept.

      You should see: A Snowpark DataFrame listing the registered feature views, plus a FeatureView object.

    3. 3.Call fs.generate_dataset(name="MY_DATASET", spine_df=..., features=[registered_fv], version="v1", spine_timestamp_col="TS").

      Why: generate_dataset always materializes an immutable, versioned snapshot.

      You should see: A Dataset whose selected version is v1.

    4. 4.Run LIST 'snow://dataset/MY_DATASET/versions/v1' in SQL.

      Why: This shows how the version is stored on disk.

      You should see: Parquet files belonging to the dataset version.

    5. 5.Call dataset.read.to_snowpark_dataframe(), then fs.retrieve_feature_values(spine_df=..., features=[registered_fv], spine_timestamp_col="TS") for a new spine.

      Why: The first call gives you training input. The second gives you scoring input from the same feature views.

      You should see: Two Snowpark DataFrames, one read from the Dataset files and one with features joined to the new spine.

    Stuck? Get a nudge

    If step 4 is denied, check that your role has USAGE or OWNERSHIP on the Dataset.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “which enriches a Snowpark DataFrame that contains the source data with the derived feature values.”
      ↩︎ Training sets from a spine DataFrame
      “Requested features are retrieved for the list of entity IDs, with point-in-time correctness with respect to the provided time stamp.”
      ↩︎ Training sets from a spine DataFrame
      “generate_dataset always materializes the result.”
      ↩︎ Snowflake Datasets: versioned, immutable snapshots
      “you can retrieve the feature view from the feature store and pass it to your model for prediction”
      ↩︎ Inference and end-to-end lineage
      “Materialized tables currently don’t guarantee immutability and have limited metadata support.”
      ↩︎ Exam trap 1
      “Training sets are ephemeral by default; they exist only as Snowpark DataFrames and are not materialized.”
      ↩︎ Prediction
      “Materialized tables currently don’t guarantee immutability and have limited metadata support.”
      ↩︎ Checkpoint
    2. 2.
      “The feature table can be used to generate training datasets used for training models in Snowpark ML, or to enrich test data”
      ↩︎ Training sets from a spine DataFrame
      “ML Lineage is automatically created when the feature store is used.”
      ↩︎ Inference and end-to-end lineage
    3. 3.
      “Each version is an immutable, point-in-time snapshot of the data managed by the Dataset.”
      ↩︎ Snowflake Datasets: versioned, immutable snapshots
      “Dataset version data is stored as evenly sized files in the Apache Parquet format.”
      ↩︎ Snowflake Datasets: versioned, immutable snapshots
      “Creating Datasets requires the CREATE DATASET schema-level privilege.”
      ↩︎ Snowflake Datasets: versioned, immutable snapshots
      “Datasets incur storage costs. Delete unused datasets to minimize costs.”
      ↩︎ Snowflake Datasets: versioned, immutable snapshots
      “automatically completing the ML Lineage graph linking source data, feature views, datasets, and models for full end-to-end governance.”
      ↩︎ Inference and end-to-end lineage
      “Reading from a Dataset requires only the USAGE privilege on the Dataset (or OWNERSHIP).”
      ↩︎ Exam trap 2

    Also cited

    Ready to test yourself?

    Practise the 15 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.