What you will be able to do
- Package a feature transformation in SQL or Snowpark as a feature view that is separate from any model code
- Choose between a Snowflake-managed and an external feature view based on who owns the pipeline
- Version a registered feature view correctly, knowing which properties you can change in place and which need a new version
- Scale feature computation by tuning the refresh warehouse, keeping refresh incremental, or moving custom work to compute pools
Key concept
Registered feature view version — A feature view packages a transformation from raw data to features as its own named, versioned object in the feature store. Once a version is registered, its pipeline definition is fixed. Model code consumes that version and never re-implements the logic.
1.Why feature logic lives outside model code
When each training script computes its own features, every model carries a private copy of the same logic, and the copies drift apart. A feature store fixes this. You define the transformation once, give it a name, and every model that needs it reads from that definition. Snowflake describes the goal as standardizing "commonly used feature transformations in a central repository, enabling reuse". That reuse is why features count as data assets and not as helper code inside a model repository.
In Snowflake these assets are ordinary database objects. A feature store is a schema. A feature view is either a dynamic table or a view, and its properties, such as name and entity, are stored as tags on that object. Because they are ordinary objects, normal role-based access control applies, and changes made through SQL appear in the Python API and the other way round. You can write the transformation in Python or SQL. The model team only has to know which feature view, and which version of it, to read.
| Feature store object | Snowflake object |
|---|---|
| feature store | schema |
| feature view | dynamic table or view |
| entity | tag |
| feature | column in a dynamic table or in a view |
Checkpoint 1 of 6· Check yourself
A platform team needs feature pipelines to fall under the same governance as other data, separate from the model repositories. What makes this possible in the Snowflake Feature Store?
A feature store is a schema and a feature view is a dynamic table or a view. Both are governed like any other Snowflake object, independently of model code.
“Feature store objects are implemented as Snowflake objects. All feature store objects are therefore subject to Snowflake access control rules.”Source: docs.snowflake.com
Sources1
2.Packaging: who runs the transformation?
The FeatureView constructor takes a Snowpark DataFrame that holds the feature logic, and you can write that logic in SQL or Snowpark Python. One parameter, refresh_freq, decides who keeps the features up to date. If you set it, the feature view is Snowflake-managed: Snowflake creates a dynamic table as the feature table and refreshes it on your schedule. The value can be a time delta, with a minimum of 1 minute, or a cron expression with a time zone.
managed_fv = FeatureView(
name="MY_MANAGED_FV",
entities=[entity],
feature_df=my_df, # Snowpark DataFrame containing feature transformations
timestamp_col="ts", # optional timestamp column name in the dataframe
refresh_freq="5 minutes", # how often feature data refreshes
desc="my managed feature view" # optional description
)If your transformations already run in a tool such as dbt, you can still publish the output as a versioned asset. Set refresh_freq to None to create an external feature view. In that case you build and maintain the feature table yourself, and the feature DataFrame is usually a plain projection of that table. External feature views are implemented as views, so they add no storage cost.
| Aspect | Snowflake-managed | External |
|---|---|---|
| refresh_freq | Time delta (min 1 minute) or cron expression | None |
| Backing object | Dynamic table | View on your feature table |
| Who runs the transformation | Snowflake, on the schedule you set | You, for example with dbt |
| Extra storage | Feature table materialized | No additional storage cost |
Checkpoint 2 of 6· Fill the gap
Your dbt project already produces the feature table. Which value registers it as an external feature view?
external_fv = FeatureView(
name="MY_EXTERNAL_FV",
entities=[entity],
feature_df=my_df, # Snowpark DataFrame referencing the feature table
timestamp_col="ts", # optional timestamp column name in the dataframe
refresh_freq= ? , # None means the feature view is external
desc="my external feature view" # optional description
)Without a refresh schedule the feature view is external. Snowflake does not refresh it, and you maintain the feature table yourself.
Source: docs.snowflake.comSources2
3.Versioning: register once, change by adding versions
Defining a feature view does not publish it. You publish it by registering it with an explicit version, and that name plus version is what training and inference code refer to.
registered_fv: FeatureView = fs.register_feature_view(
feature_view=managed_fv, # feature view created above, could also use external_fv
version="1",
block=True, # whether function call blocks until initial data is available
overwrite=False, # whether to replace existing feature view with same name/version
)After registration, the pipeline definition is immutable. That is what lets a model trained on version 1 trust that version 1 will always compute features the same way. Operational settings can still change. update_feature_view can change only three things: the refresh frequency, the warehouse where the transforms run, and the description. Any change to the logic or the columns needs a new version. Teams find registered versions with get_feature_view, list_feature_views, the Snowsight Feature Store UI, or Universal Search, and per-feature descriptions make them easier to find.
Checkpoint 3 of 6· Match them up
Match each change to how you make it on a registered feature view
Tap a term, then the definition that fits it.
Only refresh frequency, warehouse and description can be updated in place. Feature definitions and columns are fixed per version.
“The warehouse where the feature transforms execute”Source: docs.snowflake.com
Checkpoint 4 of 6· Exam question
A retail team owns a managed feature view `customer_spend` (version `v1`) that three production models already read. A data scientist needs to change the aggregation window from 30 to 90 days without disturbing those models, and the transformation logic must stay independent of any model repository. What is the correct approach?
Correct answer: A — Register the updated Snowpark transformation as a new version, `v2`, of the same feature view, and let each model move to it when it is retrained and validated.
- A. Feature view versions are immutable snapshots of a transformation. Registering `v2` beside `v1` lets existing models keep reading the old definition while new consumers adopt the new window, with no coupling to model code.
- B. `update_feature_view` only changes properties such as refresh frequency, warehouse and description. The defining query of a registered version cannot be replaced in place, and doing so would silently change features under live models anyway.
- C. Replacing the underlying dynamic table bypasses Feature Store metadata and changes values for models that were trained on the 30-day window. This produces training/serving skew rather than a controlled version change.
- D. Embedding the transformation in model code is the opposite of treating features as independent assets. Each model would reimplement the logic, and offline and online values would diverge.
Sources2
4.Scaling feature computation: warehouses, incremental refresh, compute pools
A managed feature view refreshes on a virtual warehouse. The cheapest way to scale is to keep that refresh incremental, so each run processes only new data. Incremental refresh requires change tracking on every source table. Snowflake tries to turn it on when it creates the dynamic table, but that needs OWNERSHIP of the table. If you don't own it, ask the owner to enable change tracking, or accept refresh_mode='FULL', which reads the whole source table on every refresh.
Some queries cannot be maintained incrementally. Those are fully refreshed at the specified frequency, which the docs warn "may lead to greater lag in feature refresh and higher maintenance costs." There are two remedies. You can split the query into smaller queries that do support incremental maintenance, or you can provision a larger virtual warehouse for dynamic table maintenance. Because the warehouse can be updated in place, resizing it does not need a new version.
Checkpoint 5 of 6· Fill the gap
A registered feature view's full refreshes keep falling behind schedule. Which parameter moves its transforms to a larger warehouse without creating a new version?
fs.update_feature_view(
name="<name>",
version="<version>",
refresh_freq="<new_fresh_freq>", # optional
? ="<new_warehouse>", # optional
desc="<new_description>", # optional
)The warehouse is one of the three properties you can update on a registered feature view. Scaling compute therefore leaves the feature definition and its version unchanged.
Source: docs.snowflake.comWarehouses are not the only engine. Snowflake ML's preprocessing APIs push work down to the warehouse and "distribute the processing across warehouse compute resources." For highly customizable large-scale processing, especially of unstructured data, Ray map_batches runs across nodes in a compute pool on Container Runtime. For freshness rather than throughput, stream feature views ingest events in real time and reach end-to-end freshness under two seconds. That is the near-real-time path when scheduled batch refreshes are not fresh enough.
Checkpoint 6 of 6· Exam question
A platform team keeps feature pipeline code in Git and must promote it from Dev to Test to Prod with identical transformation logic. Each environment lives in its own Snowflake database. Which promotion design best preserves consistency between environments?
Correct answer: A — Deploy the same versioned Python module from CI to each environment, passing database, schema and warehouse as parameters to `FeatureStore` and `register_feature_view`.
- A. One parameterized artifact means only environment settings differ between stages. The promoted logic is byte-identical, so a feature view tested in Test behaves the same in Prod.
- B. Independent copies drift, so what is validated in Test is not what runs in Prod. Promotion should move one tested artifact, not rely on manual reconciliation.
- C. Sharing Prod objects with lower environments prevents testing changes in isolation and exposes production data to experimentation. Dev and Test need their own feature store schemas.
- D. CSV exports lose dependency, lag and metadata details, and manual re-creation is error-prone. This is neither automated nor repeatable promotion.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.You can fix a feature's logic by updating the registered feature view in place.Why is that wrong?
update_feature_view changes only the refresh frequency, warehouse and description. Any change to the logic or columns needs a new version.
Covered in Versioning: register once, change by adding versions
2.Registering a feature view with refresh_freq=None means Snowflake will still keep it fresh on a default schedule.Why is that wrong?
With no refresh schedule the feature view is external. Snowflake does not run the pipeline, and you must maintain the feature table yourself.
Covered in Packaging: who runs the transformation?
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“A feature store lets you standardize commonly used feature transformations in a central repository, enabling reuse”
↩︎ Why feature logic lives outside model code“Changes you make via SQL are reflected in the Python API and vice versa.”
↩︎ Why feature logic lives outside model code“Feature store objects are implemented as Snowflake objects. All feature store objects are therefore subject to Snowflake access control rules.”
↩︎ Checkpoint - 2.
“Setting the refresh_freq parameter designates the feature view as Snowflake-managed.”
↩︎ Packaging: who runs the transformation?“External feature views are implemented as views on your feature table, so they incur no additional storage cost.”
↩︎ Packaging: who runs the transformation?“A feature view pipeline definition is immutable after it has been registered, providing consistent feature computation as long as the feature view exists.”
↩︎ Versioning: register once, change by adding versions“Features can also be discovered using the Snowsight Feature Store UI or Universal Search.”
↩︎ Versioning: register once, change by adding versions“provision a larger virtual warehouse for dynamic table maintenance”
↩︎ Scaling feature computation: warehouses, incremental refresh, compute pools“To qualify for incremental refresh, each source table must have change tracking enabled.”
↩︎ Scaling feature computation: warehouses, incremental refresh, compute pools“A feature view pipeline definition is immutable after it has been registered, providing consistent feature computation as long as the feature view exists.”
↩︎ Key concept“Feature definitions and columns cannot be modified. To change the features in a feature store, create a new version of the feature view.”
↩︎ Exam trap 1“You are responsible for maintaining the feature table, updating features from raw data as needed”
↩︎ Exam trap 2“Feature definitions and columns cannot be modified. To change the features in a feature store, create a new version of the feature view.”
↩︎ Prediction“The warehouse where the feature transforms execute”
↩︎ Checkpoint - 3.
“use parallel, resource-managed execution across single-node or multi-node Container Runtime environments”
↩︎ Scaling feature computation: warehouses, incremental refresh, compute pools - 4.https://docs.snowflake.com/en/developer-guide/snowflake-ml/feature-store/online-feature-storeOfficial docs
“With stream ingestion, end-to-end freshness is under 2 seconds.”
↩︎ Scaling feature computation: warehouses, incremental refresh, compute pools