What you will be able to do
- Set dynamic table target lag correctly, including DOWNSTREAM, for chained feature pipelines
- Keep offline and online feature values consistent and reason about end-to-end freshness
- Explain where tasks fit when you maintain feature tables yourself
- Promote feature definitions across dev, staging and prod with a plan-then-apply workflow
1.Refresh policies: what target lag really promises
A Snowflake-managed feature view is backed by a dynamic table, so the dynamic table's refresh policy decides how fresh your features are in every environment. The central setting is target lag.
Snowflake uses heuristics to decide when each refresh starts. Warehouse size, data volume, query complexity and pipeline depth can all push actual lag past the target. It never runs two refreshes of the same table at once, so if refreshes regularly take longer than the target lag, the data falls further behind. The minimum target lag is 60 seconds. In a chained pipeline, upstream tables refresh first, and every table reads one consistent snapshot of its inputs. Lag is measured from the base tables at the root of the pipeline.
The other form is TARGET_LAG = DOWNSTREAM. A table set this way has no schedule of its own and refreshes only when a downstream consumer needs fresh data. That makes it the right setting for intermediate feature tables. On the final table it is a trap: with no consumers, that table never refreshes, and Snowflake does not warn you.
| dt_orders | dt_orders_daily | Behavior |
|---|---|---|
| DOWNSTREAM | '10 minutes' | Recommended: dt_orders refreshes on demand whenever dt_orders_daily needs it |
| '5 minutes' | '10 minutes' | dt_orders_daily refreshes within 10 minutes, using data from dt_orders that is at most 5 minutes stale |
| '10 minutes' | DOWNSTREAM | Avoid: dt_orders_daily never refreshes |
| DOWNSTREAM | DOWNSTREAM | Neither refreshes automatically; use manual refresh |
| '1 hour' | '10 minutes' | Invalid: downstream can't be fresher than its upstream source |
Checkpoint 1 of 5· Check yourself
After promotion to prod, the final feature table in a two-table pipeline is always empty. The upstream table has TARGET_LAG = '10 minutes' and the final table has TARGET_LAG = DOWNSTREAM. What is wrong?
DOWNSTREAM takes its schedule from consumers. A table at the end of the chain has none, so it never refreshes. Put the duration on the final table and DOWNSTREAM on the upstream one.
“If a dynamic table with TARGET_LAG = DOWNSTREAM has no downstream consumers, it never refreshes automatically, and Snowflake produces no warning about this.”Source: docs.snowflake.com
Sources1
2.Offline/online parity and end-to-end freshness
Training reads the offline feature table, while inference often needs millisecond lookups. The Online Feature Store keeps the latest value per entity key in a managed Postgres serving layer. Parity comes from the store's design: it "stays integrated with your batch feature pipeline for training while serving the same feature values at inference time". You do not build separate sync jobs. You turn on online serving in the feature view's OnlineConfig, either when you create the feature view or later.
# Enable online feature serving
fs.update_feature_view(
name="<name>",
version="<version>",
online_config=OnlineConfig(
enable=True,
target_lag="10s",
store_type=OnlineStoreType.POSTGRES,
),
)Two independent settings add up here. refresh_freq controls how often the offline dynamic table refreshes. OnlineConfig.target_lag controls how often those offline values sync to the online store, anywhere from 10 seconds to 8 days. Freshness from source to the online store is therefore roughly the sum of the two. If minute-scale freshness is not enough, switch to stream feature views, which update as events arrive.
Cost also differs between environments. The online service is created once per feature store and runs continuously, so in dev or test call drop_online_service() when you are finished. Its size sets serving capacity: larger sizes handle more concurrent reads and more ingest throughput, at higher cost. An XS service cannot be resized later. If refresh fails five times in a row, the online table is suspended.
Checkpoint 2 of 5· Match them up
Match each setting to what it controls
Tap a term, then the definition that fits it.
The offline refresh and the online sync are configured separately, and freshness is roughly their sum. Stream ingestion is the path for sub-minute freshness. The service size sets serving capacity.
“the effective lag from source data to the online store is approximately refresh_freq + online_config.target_lag.”Source: docs.snowflake.com
Sources2
3.Tasks for pipelines you maintain yourself
With an external feature view, Snowflake does not refresh anything, and keeping the feature table current is your job. Snowflake tasks are the native scheduler for that work. Tasks run SQL commands and stored procedures, including Python. They can run on a schedule or be triggered by events, such as new data arriving in a stream. Task graphs can chain steps in series or in parallel.
Compute for tasks comes in two models. In the serverless model, Snowflake predicts the resources each run needs and assigns them. In the user-managed model, you run the task on a virtual warehouse that you size. Note that the sources for this lesson do not show a task example that refreshes a Feature Store table. They cover tasks generally and external feature views separately, so connecting the two is your own design choice.
Checkpoint 3 of 5· Check yourself
A team keeps an external feature view's table up to date with a stored procedure. They want Snowflake to run it on a schedule without managing a warehouse. Which option fits?
Tasks run stored procedures on a schedule, and serverless tasks size their compute automatically. Setting refresh_freq would turn the feature view into a Snowflake-managed one, and target lag applies only to dynamic tables.
“Serverless tasks: Snowflake predicts resources that are needed and assigns them automatically.”Source: docs.snowflake.com
4.Promoting feature definitions from dev to prod
Promotion is safest when the definition being promoted is a file under version control. The declarative Feature Store workflow, currently in preview and distributed as a private wheel, describes entities, data sources and feature views as YAML in a project directory. You can wire it into CI/CD to validate, test and promote features. Each deployment target names an account, database, schema and role, and manifest.yml lists one target per stage.
manifest_version: 1
type: feature_store
default_target: dev
targets:
dev:
account_identifier: MY_ORG-MY_ACCOUNT-DEV
database: MY_FS
schema: FEATURES
role: DATA_SCIENTIST
staging:
account_identifier: MY_ORG-MY_ACCOUNT-STAGING
database: MY_FS
schema: FEATURES
role: FS_ADMIN
prod:
account_identifier: MY_ORG-MY_ACCOUNT-PROD
database: MY_FS
schema: FEATURES
role: FS_ADMINSource files can leave out database and schema, because the planner fills in the resolved target. The same files therefore deploy unchanged to dev, staging and prod, which keeps definitions consistent from dev to prod. Choose the environment with --target. snow feature plan compares your local definitions with what is deployed and shows the difference. Nothing changes until snow feature apply. One constraint matters: batch feature views must be written as SQL strings and streaming ones as pandas UDFs. Snowpark DataFrame logic has no declarative equivalent and must be rewritten before you can use this workflow.
Checkpoint 4 of 5· Put it in order
Put the declarative feature development cycle in order
- 1.Plan changes to preview what would be created, updated, or deleted
- 2.Apply the plan to deploy features to a Snowflake target environment
- 3.Test by ingesting events and querying the online store
- 4.Define features as YAML and Python in a local project directory
The plan-then-apply model previews the difference against the target before anything changes, and testing comes after deployment.
“Nothing is applied until you explicitly run snow feature apply.”Source: docs.snowflake.com
Checkpoint 5 of 5· Exam question
A fraud model scores transactions through an online endpoint that must read the same feature values used during training. The team currently recomputes features in the serving application with separate Python code, and offline AUC does not match production. Which change best establishes offline/online parity?
Correct answer: A — Enable `online_config` on the managed feature view so the online store is populated from the same definition and serving reads the online values.
- A. With an online configuration, Snowflake maintains the low-latency online store from the same feature view definition that feeds offline training sets. One definition serves both paths, which removes the duplicated logic that caused the skew.
- B. Tests detect divergence only after the fact and still leave two implementations to maintain. Parity comes from sharing a single definition, not from comparing two.
- C. Changing refresh cadence does nothing to align transformation logic. It could widen staleness gaps and make the discrepancy between training and serving worse.
- D. A second copy refreshed on its own schedule is another pipeline that can drift from the original. It adds latency and maintenance instead of providing a single source of truth.
Sources5
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.TARGET_LAG = '10 minutes' means the feature table refreshes every 10 minutes and is never staler than that.Why is that wrong?
Target lag is a best-effort staleness target. Actual lag can exceed it when warehouse size, data volume or pipeline depth limit throughput.
Covered in Refresh policies: what target lag really promises
2.Setting OnlineConfig.target_lag to 10s makes online features 10 seconds behind the source.Why is that wrong?
The online sync starts from the offline table, which refreshes on its own refresh_freq. End-to-end lag is roughly the sum of the two.
3.Snowpark DataFrame feature logic can be promoted unchanged with snow feature plan and apply.Why is that wrong?
The declarative workflow accepts SQL strings for batch feature views and pandas UDFs for streaming ones. Snowpark logic must be rewritten.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The minimum target lag is 60 seconds.”
↩︎ Refresh policies: what target lag really promises“Target lag is a staleness target, not a refresh interval.”
↩︎ Refresh policies: what target lag really promises“Applications should not assume that data is always within the target lag.”
↩︎ Exam trap 1“If a dynamic table with TARGET_LAG = DOWNSTREAM has no downstream consumers, it never refreshes automatically, and Snowflake produces no warning about this.”
↩︎ Checkpoint - 2.https://docs.snowflake.com/en/developer-guide/snowflake-ml/feature-store/online-feature-storeOfficial docs
“The online store stays in sync with your offline feature pipeline automatically, so you don’t need separate sync jobs or extra infrastructure.”
↩︎ Offline/online parity and end-to-end freshness“Larger sizes serve more concurrent reads and ingestion throughput at higher cost.”
↩︎ Offline/online parity and end-to-end freshness“the effective lag from source data to the online store is approximately refresh_freq + online_config.target_lag.”
↩︎ Exam trap 2“the effective lag from source data to the online store is approximately refresh_freq + online_config.target_lag.”
↩︎ Checkpoint - 3.https://docs.snowflake.com/en/user-guide/tasks-introOfficial docs
“Tasks can run at scheduled times or can be triggered by events, such as when new data arrives in a stream.”
↩︎ Tasks for pipelines you maintain yourself“Serverless tasks: Snowflake predicts resources that are needed and assigns them automatically.”
↩︎ Checkpoint - 4.
“You are responsible for maintaining the feature table, updating features from raw data as needed”
↩︎ Tasks for pipelines you maintain yourself - 5.https://docs.snowflake.com/en/developer-guide/snowflake-ml/feature-store/feature-development-lifecycleOfficial docs
“Define multiple targets in manifest.yml to represent different stages of your workflow, such as dev, staging, prod.”
↩︎ Promoting feature definitions from dev to prod“you can integrate them into CI/CD pipelines to automatically validate, test, and promote features across development, staging, and production environments.”
↩︎ Promoting feature definitions from dev to prod“Snowpark DataFrame has no declarative equivalent.”
↩︎ Exam trap 3“Nothing is applied until you explicitly run snow feature apply.”
↩︎ Checkpoint