What you will be able to do
- Build a spine DataFrame and call generate_training_set or generate_dataset with spine_timestamp_col so feature lookups are point-in-time correct
- Explain how the ASOF join pairs a spine timestamp with a feature view's timestamp_col to stop future values leaking into training rows
- Design tests that show a training set returns no feature values from after the label event, and pick between an ephemeral, table, or Dataset output
Key concept
Point-in-time (ASOF) feature lookup — Each training row gets the feature values that existed at that row's own timestamp, not the latest values. For every spine row, the Feature Store picks the most recent feature row for that entity at or before the spine timestamp, so the model never trains on information it would not have had at prediction time.
1.The spine: where every training row gets its timestamp
Temporal leakage happens when a training row contains a feature value computed from events after the moment the label describes. A fraud model trained on a customer's transaction count *including* the transactions that came after the fraud event will look great offline and then fail in production, because at inference time those future transactions don't exist yet.
The Snowflake Feature Store prevents this through the spine. The spine is a Snowpark DataFrame that you provide. It defines which training examples exist: one row per example, holding the entity key, the time of the event, and the label. The documentation describes it this way: the spine_df "is a DataFrame containing the entity IDs in source data, the time stamp, label columns, and additional columns containing training data." You don't pull features into the spine yourself. You pass it to the feature store, and the feature store adds the feature values for you.
training_set = fs.generate_training_set(
spine_df=MySourceDataFrame,
features=[registered_fv],
save_as="data_20240101", # optional
spine_timestamp_col="TS", # optional
spine_label_cols=["LABEL1", "LABEL2"], # optional
include_feature_view_timestamp_col=False, # optional
)Notice that spine_timestamp_col is marked *optional*. That is the design decision this objective tests. The documentation's guidance is: "For time-series features, provide the timestamp column name to automate the point-in-time feature value lookup." If you supply it, "Requested features are retrieved for the list of entity IDs, with point-in-time correctness with respect to the provided time stamp." If you leave it out on time-series features, you lose the automatic point-in-time lookup, and the protection against leakage goes with it.
spine_label_cols matters for the same reason. It marks which spine columns are labels, so they stay separate from the feature columns that the lookup adds.
Checkpoint 1 of 4· Fill the gap
Which parameter turns on the automatic point-in-time lookup in this training-set call?
training_set = fs.generate_training_set(
spine_df=MySourceDataFrame,
features=[registered_fv],
save_as="data_20240101", # optional
? ="TS", # optional
spine_label_cols=["LABEL1", "LABEL2"], # optional
include_feature_view_timestamp_col=False, # optional
)spine_timestamp_col names the spine's time column, and each row's feature lookup is anchored to it. timestamp_col is the feature view's own time column, set when you define the view, not when you generate a training set.
Source: docs.snowflake.comSources1
2.How the ASOF join matches two timestamps
The lookup is an ASOF join. For each spine row, the Feature Store finds the matching entity key in the feature view and takes the most recent feature row stamped *at or before* the spine timestamp. The documentation states the result directly: the training set "reflects the features as they existed at the time of each training example, preventing future data from leaking into the model."
An ASOF join needs a timestamp on each side. The spine provides one through spine_timestamp_col. The feature view provides the other through the timestamp_col you declared when you defined it, for example timestamp_col="LAST_ACTIVITY_TS" on a batch feature view, or the optional timestamp_col on an external feature view. A feature view without a timestamp has nothing to order its rows by, so it can't take part in a time-aware match. When you design a feature view for time-series training, decide what its timestamp means (when the event happened, or when the feature was computed) before you register it.
| Timestamp | Where it is declared | What it represents |
|---|---|---|
| spine_timestamp_col | Argument to generate_training_set / generate_dataset / retrieve_feature_values | The moment each training or prediction example happened |
| timestamp_col | Argument to FeatureView when the view is defined | When each feature row's values became valid |
Some feature views require the spine timestamp. An append-only feature view keeps accumulated historical snapshots. For these, "When an append-only feature view is used as a feature source, spine_timestamp_col is required." This makes sense: a snapshot history is only useful if each training row knows which snapshot it should receive.
Checkpoint 2 of 4· Check yourself
A data scientist complains that their churn model scores 0.99 AUC offline but performs poorly in production. They built the training set with generate_training_set from a spine of customer IDs and labels, but did not pass spine_timestamp_col. What is the most likely cause?
For time-series features, the spine timestamp column is what turns on the automatic point-in-time lookup. Without it, nothing anchors each row's features to the label time, which is the classic leakage pattern.
“For time-series features, provide the timestamp column name to automate the point-in-time feature value lookup.”Source: docs.snowflake.com
Checkpoint 3 of 4· Exam question
A data scientist is building a churn model. The spine DataFrame holds `CUSTOMER_ID`, `LABEL_TS` (the date each prediction would have been made) and `CHURNED`. The `customer_activity` feature view was registered with `timestamp_col="EVENT_TS"`. Which call builds a training set that avoids temporal leakage?
Correct answer: A — Call `fs.generate_training_set(spine_df, features=[fv], spine_timestamp_col="LABEL_TS", spine_label_cols=["CHURNED"])` so each label gets feature values valid then
- A. Correct. With `spine_timestamp_col` set, the Feature Store runs an ASOF-style lookup that returns, for each spine row, the latest feature row whose timestamp is at or before the label timestamp, so no later values leak in.
- B. Incorrect. Without a spine timestamp column the lookup has no point in time to anchor on, so each label receives current feature values, which include information from after the prediction date.
- C. Incorrect. Reading the feature view returns only the latest materialized state per entity. A refresh schedule does not align values to historical label dates, so past labels get present-day features.
- D. Incorrect. The spine timestamp must be a column of the spine that states when the prediction was made. A feature timestamp is not a spine column, and comparing a column with itself provides no temporal filter.
Sources2
3.Proving no future values leaked, then freezing the result
You shouldn't trust an ASOF lookup on a new feature view until you have tested it. The Feature Store gives you a hook for this. By default, the feature view's own timestamp is left out of the output. You can bring it back: you can "include the timestamp column from the feature view by setting include_feature_view_timestamp_col". With the feature view's timestamp next to the spine timestamp in every row, two practical tests follow:
- Row-level assertion: check that no row has a feature view timestamp later than its spine timestamp. Even one such row means leakage. - Boundary probe: build a small spine with a known entity, place its timestamp just before a known feature update, and confirm you get the value from *before* the update.
Once the training set passes these tests, decide how to keep it. The documentation notes that "Training sets are ephemeral by default; they exist only as Snowpark DataFrames and are not materialized." An ephemeral set is computed again each time you use it, so it can drift as the feature view refreshes. Passing save_as writes the result to a new table, but the documentation warns that such tables don't guarantee immutability and have limited metadata support. For a frozen, reproducible training set, use generate_dataset. It takes a required name, an optional version, and the same spine arguments, and "generate_dataset always materializes the result."
| Output | How you get it | Reproducibility |
|---|---|---|
| Snowpark DataFrame | generate_training_set without save_as | Ephemeral; not materialized |
| Table | generate_training_set with save_as | Materialized, but immutability not guaranteed |
| Snowflake Dataset | generate_dataset with name and optional version | Immutable, file-based snapshot |
Checkpoint 4 of 4· Match them up
Match each option to what it does for temporal integrity
Tap a term, then the definition that fits it.
The spine timestamp drives the lookup. The feature view timestamp lets you check it. Datasets freeze the result, and save_as tables only persist it.
“Snowflake Datasets provide an immutable, file-based snapshot of data, which helps to ensure model reproducibility”Source: docs.snowflake.com
Sources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Point-in-time joins return the latest available feature value for each entity.Why is that wrong?
The ASOF join picks the most recent row at or before each spine row's timestamp. Values stamped after it are excluded, even when they are newer.
Covered in How the ASOF join matches two timestamps
2.Saving a training set with save_as gives you a frozen, reproducible snapshot.Why is that wrong?
save_as writes a table, and those tables don't guarantee immutability. generate_dataset produces an immutable, file-based Dataset that is built for reproducibility.
Covered in Proving no future values leaked, then freezing the result
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Requested features are retrieved for the list of entity IDs, with point-in-time correctness with respect to the provided time stamp.”
↩︎ The spine: where every training row gets its timestamp“the spine_df (MySourceDataFrame) is a DataFrame containing the entity IDs in source data, the time stamp, label columns, and additional columns containing training data.”
↩︎ The spine: where every training row gets its timestamp“include the timestamp column from the feature view by setting include_feature_view_timestamp_col”
↩︎ Proving no future values leaked, then freezing the result“Training sets are ephemeral by default; they exist only as Snowpark DataFrames and are not materialized.”
↩︎ Proving no future values leaked, then freezing the result“generate_dataset always materializes the result.”
↩︎ Proving no future values leaked, then freezing the result“Snowflake Datasets provide an immutable, file-based snapshot of data, which helps to ensure model reproducibility”
↩︎ Exam trap 2“For time-series features, provide the timestamp column name to automate the point-in-time feature value lookup.”
↩︎ Checkpoint“Snowflake Datasets provide an immutable, file-based snapshot of data, which helps to ensure model reproducibility”
↩︎ Checkpoint - 2.https://docs.snowflake.com/en/developer-guide/snowflake-ml/feature-store/advanced-feature-engineeringOfficial docs
“the Feature Store performs an ASOF join and selects the most recent snapshot row for each entity key at or before the spine timestamp”
↩︎ How the ASOF join matches two timestamps“When an append-only feature view is used as a feature source, spine_timestamp_col is required.”
↩︎ How the ASOF join matches two timestamps“the Feature Store performs an ASOF join and selects the most recent snapshot row for each entity key at or before the spine timestamp”
↩︎ Key concept“reflects the features as they existed at the time of each training example, preventing future data from leaking into the model”
↩︎ Exam trap 1