CertSafari
    Snowflake SnowPro Advanced: MLOps Engineer (MLA-B01)· Lessons

    Domain 1 · Lesson 3/17

    Point-in-Time Training Sets in Snowflake Feature Store

    Ensure temporal integrity and feature consistency.

    9 min read
    4% of exam
    2 sources
    Published 5 Oct 2026
    Docs as of 4 Oct 2026

    What you will be able to do

    • Build a spine DataFrame and call generate_training_set or generate_dataset with spine_timestamp_col so feature lookups are point-in-time correct
    • Explain how the ASOF join pairs a spine timestamp with a feature view's timestamp_col to stop future values leaking into training rows
    • Design tests that show a training set returns no feature values from after the label event, and pick between an ephemeral, table, or Dataset output

    Key concept

    Point-in-time (ASOF) feature lookup — Each training row gets the feature values that existed at that row's own timestamp, not the latest values. For every spine row, the Feature Store picks the most recent feature row for that entity at or before the spine timestamp, so the model never trains on information it would not have had at prediction time.

    1.The spine: where every training row gets its timestamp

    Temporal leakage happens when a training row contains a feature value computed from events after the moment the label describes. A fraud model trained on a customer's transaction count *including* the transactions that came after the fraud event will look great offline and then fail in production, because at inference time those future transactions don't exist yet.

    The Snowflake Feature Store prevents this through the spine. The spine is a Snowpark DataFrame that you provide. It defines which training examples exist: one row per example, holding the entity key, the time of the event, and the label. The documentation describes it this way: the spine_df "is a DataFrame containing the entity IDs in source data, the time stamp, label columns, and additional columns containing training data." You don't pull features into the spine yourself. You pass it to the feature store, and the feature store adds the feature values for you.

    Generating a training set from a spine. spine_timestamp_col names the column in the spine that anchors each row in time.python
    training_set = fs.generate_training_set(
        spine_df=MySourceDataFrame,
        features=[registered_fv],
        save_as="data_20240101",                    # optional
        spine_timestamp_col="TS",                   # optional
        spine_label_cols=["LABEL1", "LABEL2"],      # optional
        include_feature_view_timestamp_col=False,   # optional
    )

    Notice that spine_timestamp_col is marked *optional*. That is the design decision this objective tests. The documentation's guidance is: "For time-series features, provide the timestamp column name to automate the point-in-time feature value lookup." If you supply it, "Requested features are retrieved for the list of entity IDs, with point-in-time correctness with respect to the provided time stamp." If you leave it out on time-series features, you lose the automatic point-in-time lookup, and the protection against leakage goes with it.

    spine_label_cols matters for the same reason. It marks which spine columns are labels, so they stay separate from the feature columns that the lookup adds.

    Checkpoint 1 of 4· Fill the gap

    Which parameter turns on the automatic point-in-time lookup in this training-set call?

    training_set = fs.generate_training_set(
        spine_df=MySourceDataFrame,
        features=[registered_fv],
        save_as="data_20240101",                    # optional
         ? ="TS",                   # optional
        spine_label_cols=["LABEL1", "LABEL2"],      # optional
        include_feature_view_timestamp_col=False,   # optional
    )

    Sources1

    2.How the ASOF join matches two timestamps

    The lookup is an ASOF join. For each spine row, the Feature Store finds the matching entity key in the feature view and takes the most recent feature row stamped *at or before* the spine timestamp. The documentation states the result directly: the training set "reflects the features as they existed at the time of each training example, preventing future data from leaking into the model."

    An ASOF join needs a timestamp on each side. The spine provides one through spine_timestamp_col. The feature view provides the other through the timestamp_col you declared when you defined it, for example timestamp_col="LAST_ACTIVITY_TS" on a batch feature view, or the optional timestamp_col on an external feature view. A feature view without a timestamp has nothing to order its rows by, so it can't take part in a time-aware match. When you design a feature view for time-series training, decide what its timestamp means (when the event happened, or when the feature was computed) before you register it.

    The two timestamps an ASOF lookup compares
    TimestampWhere it is declaredWhat it represents
    spine_timestamp_colArgument to generate_training_set / generate_dataset / retrieve_feature_valuesThe moment each training or prediction example happened
    timestamp_colArgument to FeatureView when the view is definedWhen each feature row's values became valid

    Some feature views require the spine timestamp. An append-only feature view keeps accumulated historical snapshots. For these, "When an append-only feature view is used as a feature source, spine_timestamp_col is required." This makes sense: a snapshot history is only useful if each training row knows which snapshot it should receive.

    Checkpoint 2 of 4· Check yourself

    A data scientist complains that their churn model scores 0.99 AUC offline but performs poorly in production. They built the training set with generate_training_set from a spine of customer IDs and labels, but did not pass spine_timestamp_col. What is the most likely cause?

    Checkpoint 3 of 4· Exam question

    A data scientist is building a churn model. The spine DataFrame holds `CUSTOMER_ID`, `LABEL_TS` (the date each prediction would have been made) and `CHURNED`. The `customer_activity` feature view was registered with `timestamp_col="EVENT_TS"`. Which call builds a training set that avoids temporal leakage?

    Sources2

    3.Proving no future values leaked, then freezing the result

    You shouldn't trust an ASOF lookup on a new feature view until you have tested it. The Feature Store gives you a hook for this. By default, the feature view's own timestamp is left out of the output. You can bring it back: you can "include the timestamp column from the feature view by setting include_feature_view_timestamp_col". With the feature view's timestamp next to the spine timestamp in every row, two practical tests follow:

    - Row-level assertion: check that no row has a feature view timestamp later than its spine timestamp. Even one such row means leakage. - Boundary probe: build a small spine with a known entity, place its timestamp just before a known feature update, and confirm you get the value from *before* the update.

    Once the training set passes these tests, decide how to keep it. The documentation notes that "Training sets are ephemeral by default; they exist only as Snowpark DataFrames and are not materialized." An ephemeral set is computed again each time you use it, so it can drift as the feature view refreshes. Passing save_as writes the result to a new table, but the documentation warns that such tables don't guarantee immutability and have limited metadata support. For a frozen, reproducible training set, use generate_dataset. It takes a required name, an optional version, and the same spine arguments, and "generate_dataset always materializes the result."

    Three ways to hold a point-in-time training set
    OutputHow you get itReproducibility
    Snowpark DataFramegenerate_training_set without save_asEphemeral; not materialized
    Tablegenerate_training_set with save_asMaterialized, but immutability not guaranteed
    Snowflake Datasetgenerate_dataset with name and optional versionImmutable, file-based snapshot

    Checkpoint 4 of 4· Match them up

    Match each option to what it does for temporal integrity

    Tap a term, then the definition that fits it.

    Sources1

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Point-in-time joins return the latest available feature value for each entity.Why is that wrong?

      The ASOF join picks the most recent row at or before each spine row's timestamp. Values stamped after it are excluded, even when they are newer.

      Covered in How the ASOF join matches two timestamps

    2. 2.Saving a training set with save_as gives you a frozen, reproducible snapshot.Why is that wrong?

      save_as writes a table, and those tables don't guarantee immutability. generate_dataset produces an immutable, file-based Dataset that is built for reproducibility.

      Covered in Proving no future values leaked, then freezing the result

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Requested features are retrieved for the list of entity IDs, with point-in-time correctness with respect to the provided time stamp.”
      ↩︎ The spine: where every training row gets its timestamp
      “the spine_df (MySourceDataFrame) is a DataFrame containing the entity IDs in source data, the time stamp, label columns, and additional columns containing training data.”
      ↩︎ The spine: where every training row gets its timestamp
      “include the timestamp column from the feature view by setting include_feature_view_timestamp_col”
      ↩︎ Proving no future values leaked, then freezing the result
      “Training sets are ephemeral by default; they exist only as Snowpark DataFrames and are not materialized.”
      ↩︎ Proving no future values leaked, then freezing the result
      “generate_dataset always materializes the result.”
      ↩︎ Proving no future values leaked, then freezing the result
      “Snowflake Datasets provide an immutable, file-based snapshot of data, which helps to ensure model reproducibility”
      ↩︎ Exam trap 2
      “For time-series features, provide the timestamp column name to automate the point-in-time feature value lookup.”
      ↩︎ Checkpoint
      “Snowflake Datasets provide an immutable, file-based snapshot of data, which helps to ensure model reproducibility”
      ↩︎ Checkpoint
    2. 2.
      “the Feature Store performs an ASOF join and selects the most recent snapshot row for each entity key at or before the spine timestamp”
      ↩︎ How the ASOF join matches two timestamps
      “When an append-only feature view is used as a feature source, spine_timestamp_col is required.”
      ↩︎ How the ASOF join matches two timestamps
      “the Feature Store performs an ASOF join and selects the most recent snapshot row for each entity key at or before the spine timestamp”
      ↩︎ Key concept
      “reflects the features as they existed at the time of each training example, preventing future data from leaking into the model”
      ↩︎ Exam trap 1

    Continue to page 2 of 2

    Offline/Online and Dev/Prod Feature Consistency

    Spotted a mistake, or was something unclear? Tell us.