What you will be able to do
- Create or connect to a feature store with the FeatureStore constructor and choose the right CreationMode
- Explain how feature store objects map to schemas, dynamic tables, views, tags and columns
- Define and register entities and know which of their properties are immutable
- Build Snowflake-managed and external feature views from Snowpark DataFrames and say when each is backed by a dynamic table
- Register, version, update and discover feature views as shared metadata
- Explain how a central feature store enables feature reuse and a single source of truth across teams
- Describe how registered feature views act as centralized metadata objects that other teams and use cases can discover and reuse
Key concept
Feature store objects are Snowflake objects — A feature store is just a schema. Its feature views are dynamic tables or views, its entities are tags, and its features are columns. That means SQL and the Python API work on the same objects, and normal role-based access control applies to all of them.
1.A feature store is a schema
A feature store gives teams one central place to define common feature transformations. They can then reuse those transformations instead of copying them. It also keeps features current as new source data arrives, so everyone reads from a single source of truth. In Snowflake there is no separate service behind this. The documentation states it directly: "A feature store in Snowflake is simply a schema." You can create a new schema for it or use an existing one.
Everything else in the feature store sits on ordinary Snowflake objects. A feature view is backed by either a dynamic table or a view. An entity is a tag, and each feature is a column. Properties of a feature view, such as its name and entity, are also stored as tags on the underlying dynamic table or view. Because these are real objects, you can query or change them with SQL, and those changes show up in the Python API (and the other way round). To delete a feature store completely, first make sure the schema holds nothing else, then drop the schema.
Checkpoint 1 of 7· Match them up
Match each Feature Store concept to the Snowflake object that implements it
Tap a term, then the definition that fits it.
The overview maps each feature store object to a native Snowflake object. That is why SQL changes and Python API changes stay in sync.
“feature view | dynamic table or view”Source: docs.snowflake.com
You create a feature store, or connect to one, with the FeatureStore class. It is part of the snowflake-ml-python package. The constructor takes a Snowpark session, a database name, a feature store name and a default warehouse. The creation_mode parameter decides what happens if the schema does not exist yet. The default, CreationMode.FAIL_IF_NOT_EXIST, raises an exception. CreationMode.CREATE_IF_NOT_EXIST creates the schema. Normally an administrator role creates the feature store schema and its roles. Everyone else then connects using the default mode. The documentation also suggests keeping feature stores in a dedicated database because that makes them easier to replicate. A user with access to several feature stores can combine feature views from all of them when building training and inference datasets.
| Mode | Behaviour |
|---|---|
| CreationMode.FAIL_IF_NOT_EXIST | Throws an exception if the feature store does not exist. This is the default. |
| CreationMode.CREATE_IF_NOT_EXIST | Creates the feature store schema in the given database if it does not exist. |
Checkpoint 2 of 7· Fill the gap
An administrator is setting up a brand-new feature store. Which value completes the call?
from snowflake.ml.feature_store import FeatureStore, CreationMode
fs = FeatureStore(
session=session,
database="MY_DB",
name="MY_FEATURE_STORE",
default_warehouse="MY_WH",
creation_mode=CreationMode. ? ,
)Only CREATE_IF_NOT_EXIST creates the schema. The default, FAIL_IF_NOT_EXIST, is the mode for connecting to a store that already exists, and CREATE_OR_REPLACE is not a mode.
Source: docs.snowflake.comCheckpoint 3 of 7· Exam question
A retail team has a customer-level purchase aggregation in a Snowpark DataFrame keyed on `CUSTOMER_ID`. Several feature views will be built over different source tables, and all must be retrievable for the same business object. What is the correct way to model this in the Feature Store?
Correct answer: A — Register one `Entity` named `CUSTOMER` with join_keys `["CUSTOMER_ID"]`, then pass it as `entities` on every feature view whose DataFrame has that column.
- A. Correct. An Entity is a registered, tag-backed definition whose join keys must exist as columns in each feature view's DataFrame, which lets views and spines join consistently.
- B. Incorrect. Entities are designed to be shared; many feature views can reference one entity, and duplicating them defeats centralized reuse and consistent joins.
- C. Incorrect. Join keys are declared explicitly on the Entity object; the store does not infer them from tags you place on source tables.
- D. Incorrect. Join keys identify the subject of the features, such as a customer, never the prediction label; using a label as a key would leak the target.
The architecture exists to give you centralized feature management and reuse. The central repository standardizes commonly used feature transformations, so teams reuse one definition instead of rebuilding it. That reduces duplicated data and effort and improves productivity. Features are updated on new source data, so the store is a single source of truth that is correct, consistent and fresh. Extracting features from raw data in one consistent way also makes production ML pipelines more robust.
The central store is hosted natively inside Snowflake. Your data stays under your control and governance and never leaves Snowflake. Access is managed with fine-grained role-based access control, because every feature store object is a Snowflake object. The Snowsight Feature Store UI lets people search for and discover features. In practice, one administrator-created schema holds the shared entities and feature views, and many teams and models read from it. A user with access to more than one store can still combine feature views from several of them.
2.Entities: the data model and its join keys
Feature views are grouped by entity. An entity is the subject a feature describes. For example, a streaming service might have users and movies. Entities do two jobs. They organise feature views so people can find them, and they hold the names of the key columns used to join features back to other data. As the docs put it, entities store the names of the key columns you can use to join the extracted features back to the original data. A feature view can be tagged with more than one entity. For example, a feature about a user's activity on a particular movie relates to both.
You define an entity with a name, a list of join_keys and an optional description, then register it in the store:
from snowflake.ml.feature_store import Entity
entity = Entity(
name="MY_ENTITY",
join_keys=["UNIQUE_ID"],
desc="my entity"
)
fs.register_entity(entity)Several other methods manage entities. list_entities() returns the registered entities as a Snowpark DataFrame. get_entity(name=...) returns one entity, for example so you can read its join_keys. delete_entity removes one. The rules after registration are strict. update_entity can change only the description. Join keys and the other properties cannot be changed, so you create a new entity if you need different keys. You also cannot delete an entity while any feature view references it. Because entities are tags, they count toward the limit of 10,000 tags per account and 50 unique tags per object.
Checkpoint 4 of 7· Check yourself
A team registered a CUSTOMER entity with join key CUST_ID. The upstream system now uses CUSTOMER_KEY instead. What should they do?
update_entity can change only the description, and join keys are immutable. An entity referenced by feature views cannot be deleted, so the supported route is a new entity.
“Other aspects of the entity, such as its join keys, are immutable. To change these, create a new entity.”Source: docs.snowflake.com
Sources3
3.Feature views: managed (dynamic tables) or external
A feature view wraps the transformation of raw data into one or more related features. All features in one view refresh on the same schedule. In Python you build one with snowflake.ml.feature_store.FeatureView, and its constructor accepts a Snowpark DataFrame that contains the feature generation logic. That DataFrame is how you build a feature store with Snowpark. It must include the join_keys columns of the associated entities. It also needs a timestamp column name if the view contains time-series features. You can write the transformations in SQL or in Snowpark Python.
Setting refresh_freq makes the view Snowflake-managed. The value can be a time delta (minimum 1 minute) or a cron expression with a time zone. This is where dynamic tables come in: a Snowflake-managed feature view uses a dynamic table as the feature table. Snowflake refreshes it on your schedule and handles new data incrementally where it can.
managed_fv = FeatureView(
name="MY_MANAGED_FV",
entities=[entity],
feature_df=my_df, # Snowpark DataFrame containing feature transformations
timestamp_col="ts", # optional timestamp column name in the dataframe
refresh_freq="5 minutes", # how often feature data refreshes
desc="my managed feature view" # optional description
)Incremental refresh has two conditions. First, each source table must have change tracking enabled. If it is not enabled, Snowflake tries to turn it on when it creates the dynamic table, and that requires OWNERSHIP of the table. If you don't own the table, you can ask the owner to enable change tracking, or you can create the view with refresh_mode='FULL', which fully reads the source table for each refresh. Second, the query itself must support incremental maintenance. If it doesn't, the table will be fully refreshed from the query at the specified frequency. That can mean more refresh lag and higher maintenance costs. The documented fixes are to split the query into smaller queries that do support incremental maintenance, or to use a larger warehouse.
An external feature view is one where refresh_freq=None. It is meant for features that are already computed outside the Feature Store, for example by dbt. You create and maintain the feature table yourself. The feature DataFrame reads from that feature table, not from raw data, and is usually a simple projection with no transformations. External feature views are implemented as views on your feature table, so they add no storage cost. They still get the same entities, versioning and discovery as managed views, which keeps outside pipelines consistent with the rest of the store.
| Aspect | Snowflake-managed | External |
|---|---|---|
| How it is declared | refresh_freq set to a time delta or cron expression | refresh_freq=None |
| Backing Snowflake object | Dynamic table as the feature table | View on a feature table you maintain |
| Who keeps features fresh | Snowflake, incrementally where the query allows | You, e.g. with dbt |
| What feature_df contains | Transformation logic over source data | Usually a simple projection of the feature table |
| Extra storage | The dynamic table stores the feature values | None (implemented as a view) |
Checkpoint 5 of 7· Check yourself
A managed feature view's DataFrame aggregates a large orders table, and the query cannot be maintained incrementally by a dynamic table. What happens?
Snowflake falls back to full refresh on the same schedule. To fix it, break the query into incrementally maintainable parts or use a larger warehouse.
“the table will be fully refreshed from the query at the specified frequency.”Source: docs.snowflake.com
Checkpoint 6 of 7· Exam question
A data engineering team already maintains a curated feature table in Snowflake using dbt on its own schedule. The ML team wants those features discoverable in the Feature Store without Snowflake refreshing or duplicating the data. Which configuration achieves this?
Correct answer: B — Register a `FeatureView` over the existing table with `refresh_freq=None`, so it is an external feature view backed by a view and refreshed only by dbt.
- A. Incorrect. Cloning adds an unnecessary copy and schedule, and a managed view over the clone still builds a dynamic table when the original data can be referenced in place.
- B. Correct. Setting refresh_freq to None makes the feature view external: it is a view over data maintained elsewhere, so there is no dynamic table, no refresh compute and no duplicated storage.
- C. Incorrect. A refresh_freq creates a managed feature view backed by a dynamic table, which would copy the data and spend refresh compute on top of what dbt already does.
- D. Incorrect. Entities describe the subject of features, not tables, and converting a dbt-owned table into a dynamic table changes who controls its refresh and breaks the dbt pipeline.
Sources4
4.Registering, versioning and discovering features
Defining a FeatureView object does not share it with anyone. You do that by registering it with register_feature_view and giving it a version. Both managed and external views are registered this way. After registration, managed views start their incremental maintenance and automatic refresh. block=True makes the call wait until the initial data is available. overwrite=False stops you from replacing an existing view with the same name and version.
registered_fv: FeatureView = fs.register_feature_view(
feature_view=managed_fv, # feature view created above, could also use external_fv
version="1",
block=True, # whether function call blocks until initial data is available
overwrite=False, # whether to replace existing feature view with same name/version
)Versioning works because a registered pipeline definition is immutable, so features are computed the same way for as long as that version exists. update_feature_view can change only three properties: the refresh frequency, the warehouse where the transforms run, and the description. Feature definitions and columns can't be changed. To change features, you register a new version of the feature view. Consumers keep using version 1 until they choose to move.
Registration also makes features searchable shared metadata. attach_feature_desc adds a short description to each feature, which helps people find features in Snowsight Universal Search. get_feature_view(name, version) fetches a specific version. list_feature_views filters by entity or name and returns the results as a Snowpark DataFrame. People can also browse the Snowsight Feature Store UI. For teams that keep definitions in source control, a declarative snow feature plan/apply CLI describes entities and feature views as YAML files that can go through CI/CD. It is in preview and distributed as a private wheel.
Checkpoint 7 of 7· Check yourself
A registered feature view needs one more column and should refresh hourly instead of every 5 minutes. Which approach is supported?
update_feature_view can change only the refresh frequency, warehouse and description. Adding a feature requires a new version of the feature view.
“Feature definitions and columns cannot be modified. To change the features in a feature store, create a new version of the feature view.”Source: docs.snowflake.com
Registering turns a feature definition into a centralized metadata object, which is what lets different use cases and teams discover and reuse it. A registered feature view is a Snowflake object in the feature store schema. Its properties, such as name and entity, are implemented as tags on the dynamic table or view. Because the feature view is tagged with its entity, any team can call list_feature_views with an entity name to find the existing views for, say, users, before building a new one. Per-feature descriptions, the Snowsight Feature Store UI and Universal Search help people locate features without asking the team that built them.
Reuse works across models and teams because each consumer refers to a view by name and version. Different models can read the same registered view, and every one of them gets the same computation. Since the registered objects are ordinary Snowflake objects, access control decides who can see and use them. A team that wants a changed feature registers a new version rather than altering the shared one.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.An entity's join keys can be corrected later with update_entity.Why is that wrong?
update_entity changes only the description. Join keys are immutable, so a different key needs a new entity.
Covered in Entities: the data model and its join keys
2.Every feature view is backed by a dynamic table.Why is that wrong?
Only Snowflake-managed views use a dynamic table. External views (refresh_freq=None) are views on a feature table you maintain, and they add no storage cost.
Covered in Feature views: managed (dynamic tables) or external
3.update_feature_view can change a registered view's feature logic.Why is that wrong?
Only refresh frequency, warehouse and description can be updated. The pipeline definition is immutable after registration, so feature changes need a new version.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“A feature store in Snowflake is simply a schema.”
↩︎ A feature store is a schema“Users who have access to more than one feature store can combine feature views from multiple feature stores”
↩︎ A feature store is a schema“standardize commonly used feature transformations in a central repository, enabling reuse, helping to reduce duplication of data and effort”
↩︎ A feature store is a schema“Your data remains secure, completely under your control and governance, and never leaves Snowflake.”
↩︎ A feature store is a schema“Properties of feature views (such as name and entity) are implemented as tags on dynamic tables or views.”
↩︎ Registering, versioning and discovering features“Feature store objects are implemented as Snowflake objects. All feature store objects are therefore subject to Snowflake access control rules.”
↩︎ Key concept“feature view | dynamic table or view”
↩︎ Checkpoint - 2.
“Creating a feature store creates a schema in the specified database with the specified feature store name.”
↩︎ A feature store is a schema - 3.
“entities store the names of the key columns you can use to join the extracted features back to the original data.”
↩︎ Entities: the data model and its join keys“Entities that are referenced by any feature view cannot be deleted.”
↩︎ Entities: the data model and its join keys“Entities are implemented as tags and are subject to the limit of 10,000 tags per account and 50 unique tags per object.”
↩︎ Entities: the data model and its join keys“Other aspects of the entity, such as its join keys, are immutable. To change these, create a new entity.”
↩︎ Exam trap 1“Other aspects of the entity, such as its join keys, are immutable. To change these, create a new entity.”
↩︎ Checkpoint - 4.
“The FeatureView constructor accepts a Snowpark DataFrame that contains the feature generation logic.”
↩︎ Feature views: managed (dynamic tables) or external“A Snowflake-managed feature view uses a dynamic table as the feature table.”
↩︎ Feature views: managed (dynamic tables) or external“To qualify for incremental refresh, each source table must have change tracking enabled.”
↩︎ Feature views: managed (dynamic tables) or external“This may lead to greater lag in feature refresh and higher maintenance costs.”
↩︎ Feature views: managed (dynamic tables) or external“The feature DataFrame is based on the feature table, not on the raw data source”
↩︎ Feature views: managed (dynamic tables) or external“Adding per-feature descriptions to the FeatureView makes it easier to find features using Snowsight Universal Search.”
↩︎ Registering, versioning and discovering features“Features can also be discovered using the Snowsight Feature Store UI or Universal Search.”
↩︎ Registering, versioning and discovering features“External feature views are implemented as views on your feature table, so they incur no additional storage cost.”
↩︎ Exam trap 2“A feature view pipeline definition is immutable after it has been registered”
↩︎ Exam trap 3“A feature view is considered Snowflake-managed if you provide a schedule for refreshing it.”
↩︎ Prediction“the table will be fully refreshed from the query at the specified frequency.”
↩︎ Checkpoint“Feature definitions and columns cannot be modified. To change the features in a feature store, create a new version of the feature view.”
↩︎ Checkpoint - 5.https://docs.snowflake.com/en/developer-guide/snowflake-ml/feature-store/feature-development-lifecycleOfficial docs
“you can integrate them into CI/CD pipelines to automatically validate, test, and promote features”
↩︎ Registering, versioning and discovering features