What you will be able to do
- Explain how refresh_freq and OnlineConfig.target_lag together determine how far the online store lags behind source data
- Verify that offline training features match online serving features, using feature groups and the session timezone fix for offline reads
- Keep feature computation identical across Dev and Prod using immutable, versioned feature views and declarative deployment targets
1.One pipeline feeding two stores
A model is trained on offline features and served from online features. If the two disagree, you get training-serving skew: the model sees different inputs in production from the ones it learned on. Snowflake's answer is to avoid building two pipelines at all. "Online feature serving stores the latest feature values keyed by entity," and "The online store stays in sync with your offline feature pipeline automatically". You don't write separate sync jobs.
For batch feature views, two independent settings control this: "Feature values computed in the offline pipeline are synchronized to the online store on a schedule you configure with target_lag." Before that sync, the offline dynamic table refreshes from source on the feature view's refresh_freq.
| Setting | Where it is set | What it controls |
|---|---|---|
| refresh_freq | FeatureView definition | How often the offline dynamic table refreshes from source data |
| target_lag | OnlineConfig | How often offline values sync to the online store (10 seconds to 8 days) |
| store_type | OnlineConfig | Online store implementation; OnlineStoreType.POSTGRES for the Online Feature Store, default OnlineStoreType.HYBRID_TABLE |
Because "These settings act independently," an online value can be older than either setting alone suggests: "the effective lag from source data to the online store is approximately refresh_freq + online_config.target_lag." When you check parity, an online value that trails the offline table by up to target_lag is expected, not a defect.
Operational controls keep both stores in step. Use suspend_feature_view and resume_feature_view to pause and restart refresh: "These calls affect both the offline dynamic table and the online store so offline and online data stay consistent." Use get_refresh_history with StoreType.OFFLINE or StoreType.ONLINE to see which side last refreshed and when.
Checkpoint 1 of 5· Fill the gap
Which OnlineConfig parameter sets how often offline values sync to the online store?
fs.update_feature_view(
name="<name>",
version="<version>",
online_config=OnlineConfig(
enable=True,
? ="10s",
store_type=OnlineStoreType.POSTGRES,
),
)target_lag inside OnlineConfig controls the offline-to-online sync. refresh_freq is a separate FeatureView setting that controls the offline dynamic table refresh.
Source: docs.snowflake.comSources1
2.Verifying that offline training matches online serving
The strongest parity guarantee is structural: train on exactly the columns you will serve. If a model consumes a bundle of features in production, pass a feature group to generate_training_set instead of a list of feature views. "This reproduces the same columns offline that read_feature_group returns online, reducing training-serving skew." There is a precondition: "All source feature views in the group must have online serving enabled with store_type=OnlineStoreType.POSTGRES."
fg = fs.get_feature_group("USER_FG", "v1")
training_set = fs.generate_training_set(
spine_df=MySourceDataFrame,
feature_group=fg,
spine_timestamp_col="TS", # optional
spine_label_cols=["LABEL1", "LABEL2"], # optional
)To check values, compare a read with store_type="online" against a read with store_type="offline" for the same entity keys. Know the documented trap first: "The same entity keys can return correct values from an online read (store_type="online") but incorrect values from an offline read". This "affects stream feature views synced to the offline store and other offline feature views whose timestamp columns use TIMESTAMP_NTZ (UTC) or TIMESTAMP_LTZ." Postgres online stores treat timestamps as UTC, and the SDK doesn't report which timezone applies on read. The fix is to set the session timezone to UTC before reading offline:
ALTER SESSION SET TIMEZONE = 'UTC';
One more source of apparent mismatch applies to hybrid-table online storage. If the source has duplicate keys, "When multiple rows with the same primary key are found in the source, Snowflake only ingests the version with the most recent timestamp". Without a timestamp_col, the most recently processed row wins instead. That is a weaker basis for a deterministic comparison.
Checkpoint 2 of 5· Check yourself
A team wants the strongest structural guarantee that its training columns match what its online path returns at inference. Which approach do the docs recommend?
Passing a feature group to generate_training_set reproduces offline exactly the columns that read_feature_group returns online, which is the documented way to reduce training-serving skew.
“This reproduces the same columns offline that read_feature_group returns online, reducing training-serving skew.”Source: docs.snowflake.com
Checkpoint 3 of 5· Exam question
Which mechanism does the Snowflake Feature Store use when `generate_training_set` receives a `spine_timestamp_col` and feature views that define a `timestamp_col`?
Correct answer: D — An `ASOF JOIN` that matches each spine row to the latest feature row for the same entity key at or before the spine timestamp
- A. Incorrect. An equality join would drop most spine rows, because label times rarely match feature timestamps exactly. Point-in-time retrieval needs a nearest-prior match instead.
- B. Incorrect. Time Travel is limited by retention and operates on one timestamp per query, not per spine row. The Feature Store does not rely on it for point-in-time lookups.
- C. Incorrect. A cross join would be prohibitively expensive at training-set scale. Snowflake provides a purpose-built temporal join that avoids materializing that product.
- D. Correct. The point-in-time lookup is implemented with an ASOF JOIN whose match condition keeps the closest feature row not later than the spine time, per entity key.
3.Keeping Dev and Prod features identical
Dev/Prod consistency means the feature a model was validated on in development is computed the same way in production. Two Feature Store properties support this.
First, registered definitions can't drift. "A feature view pipeline definition is immutable after it has been registered, providing consistent feature computation as long as the feature view exists." update_feature_view can change only the refresh frequency, the warehouse, the description, and online config. "Feature definitions and columns cannot be modified. To change the features in a feature store, create a new version of the feature view." So a name plus a version identifies one fixed computation. Promoting a model means promoting it together with the exact feature view versions it was trained on.
Second, the same definition files can be deployed to every environment. The declarative workflow (a preview CLI distributed as a private wheel) describes entities, data sources, and feature views as YAML in a project directory. You "Define multiple targets in manifest.yml to represent different stages of your workflow, such as dev, staging, prod." Source files can leave out database and schema, "so the same files work against any target listed in manifest.yml."
default_target: dev
targets:
dev:
account_identifier: MY_ORG-MY_ACCOUNT-DEV
database: MY_FS
schema: FEATURES
role: DATA_SCIENTIST
staging:
account_identifier: MY_ORG-MY_ACCOUNT-STAGING
database: MY_FS
schema: FEATURES
role: FS_ADMIN
prod:
account_identifier: MY_ORG-MY_ACCOUNT-PROD
database: MY_FS
schema: FEATURES
role: FS_ADMINDetecting drift before promotion is what snow feature plan does: "The planner compares your local project definitions against the deployed state in the target environment and shows a diff". "Nothing is applied until you explicitly run snow feature apply." Running plan --target prod against the same files that are already in dev shows any divergence. snow feature describe shows the deployed state of an entity or feature view. After applying, "you can validate your deployment by ingesting test events and querying the online store directly from the CLI". Because definitions are plain files, "you can integrate them into CI/CD pipelines to automatically validate, test, and promote features across development, staging, and production environments."
One catch for teams that prototyped imperatively: "Snowpark DataFrame is not currently supported for defining or deploying features with snow feature." Batch logic has to be rewritten as a SQL string, and streaming logic as a pandas UDF. After a rewrite, re-run the parity checks on the new definition.
Checkpoint 4 of 5· Put it in order
Put the declarative feature development cycle in order
- 1.Iterate by editing definitions and repeating the plan-apply loop
- 2.Run snow feature apply to deploy to the target environment
- 3.Run snow feature plan to preview what would be created, updated, or deleted
- 4.Define features as YAML and Python in a local project directory
- 5.Test by ingesting events and querying the online store
Plan always comes before apply, so changes are reviewed as a diff before anything reaches the target. Testing happens against the deployed state.
“Nothing is applied until you explicitly run snow feature apply.”Source: docs.snowflake.com
Checkpoint 5 of 5· Exam question
An auditor asks a team to retrain a fraud model three months from now on exactly the same training rows. The feature views are managed and keep refreshing from new source data. What should the team do when producing the training data today?
Correct answer: B — Call `fs.generate_dataset` with a name and version so the result is materialized as an immutable, versioned Snowflake Dataset
- A. Incorrect. The query is only a definition. Managed feature views are refreshed in place, so rerunning later reads different underlying data, and restated or backfilled rows can change results.
- B. Correct. A Dataset stores a materialized snapshot of the training data under a version, so the exact rows can be reloaded later regardless of what the feature views have since refreshed.
- C. Incorrect. Time Travel retention is bounded, costs storage, and reproduces a table state, not the specific training join. It is not designed as a reproducibility record for training data.
- D. Incorrect. Feature view versions identify feature definitions, not data snapshots. Creating one per refresh multiplies objects and still leaves each version's data changing as it refreshes.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.OnlineConfig.target_lag alone sets how stale online features can be relative to source data.Why is that wrong?
target_lag covers only the offline-to-online sync. The offline dynamic table refreshes separately on refresh_freq, so the end-to-end lag is roughly the sum of the two.
Covered in One pipeline feeding two stores
2.If offline and online reads disagree for the same keys, the online sync is broken.Why is that wrong?
A documented limitation lets offline reads return wrong values that depend on the session timezone while the online read is correct. Set the session timezone to UTC before reading offline.
Covered in Verifying that offline training matches online serving
3.To fix a feature in production, update the registered feature view's definition in place.Why is that wrong?
Registered definitions are immutable, and update_feature_view can't change features or columns. Changing a feature means registering a new version, which keeps each version's computation fixed.
Covered in Keeping Dev and Prod features identical
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.snowflake.com/en/developer-guide/snowflake-ml/feature-store/online-feature-storeOfficial docs
“Feature values computed in the offline pipeline are synchronized to the online store on a schedule you configure with target_lag.”
↩︎ One pipeline feeding two stores“These calls affect both the offline dynamic table and the online store so offline and online data stay consistent.”
↩︎ One pipeline feeding two stores“When multiple rows with the same primary key are found in the source, Snowflake only ingests the version with the most recent timestamp.”
↩︎ Verifying that offline training matches online serving“the effective lag from source data to the online store is approximately refresh_freq + online_config.target_lag.”
↩︎ Exam trap 1 - 2.
“All source feature views in the group must have online serving enabled with store_type=OnlineStoreType.POSTGRES.”
↩︎ Verifying that offline training matches online serving“This reproduces the same columns offline that read_feature_group returns online, reducing training-serving skew.”
↩︎ Checkpoint - 3.
“This affects stream feature views synced to the offline store and other offline feature views whose timestamp columns use TIMESTAMP_NTZ (UTC) or TIMESTAMP_LTZ.”
↩︎ Verifying that offline training matches online serving“A feature view pipeline definition is immutable after it has been registered, providing consistent feature computation as long as the feature view exists.”
↩︎ Keeping Dev and Prod features identical“The same entity keys can return correct values from an online read (store_type="online") but incorrect values from an offline read”
↩︎ Exam trap 2“Feature definitions and columns cannot be modified. To change the features in a feature store, create a new version of the feature view.”
↩︎ Exam trap 3“When you call read_feature_view with store_type="offline", feature values can depend on the session TIMEZONE parameter.”
↩︎ Prediction - 4.https://docs.snowflake.com/en/developer-guide/snowflake-ml/feature-store/feature-development-lifecycleOfficial docs
“The planner compares your local project definitions against the deployed state in the target environment and shows a diff”
↩︎ Keeping Dev and Prod features identical“Snowpark DataFrame is not currently supported for defining or deploying features with snow feature.”
↩︎ Keeping Dev and Prod features identical“Nothing is applied until you explicitly run snow feature apply.”
↩︎ Checkpoint