What you will be able to do
- Map Feature Store objects (feature store, entity, feature view, feature) to the Snowflake objects behind them
- Create managed and external feature views and explain how refresh_freq decides which one you get
- Register, version, retrieve and update feature views without disturbing existing consumers
- Generate point-in-time correct training data with generate_dataset and spine_timestamp_col
1.What a feature store is in Snowflake
Feature engineering means deciding which features a model needs and how to derive them from raw data. The Snowflake Feature Store keeps those derivations in one central place, so teams reuse a single definition instead of each re-implementing, say, a 30-day average spend. It keeps the features current as new source data arrives, supports transformations written in Python or SQL, and provides point-in-time correct backfill through ASOF JOIN.
The Feature Store adds no separate storage system. Each of its objects is an ordinary Snowflake object, so normal role-based access control applies, and changes you make in SQL show up in the Python API (and the reverse). A feature view packages a Python or SQL pipeline that produces related features, all refreshed together. An entity is the subject the features describe, such as a user or a movie, and feature views are organised by entity.
| Feature store object | Snowflake object |
|---|---|
| feature store | schema |
| feature view | dynamic table or view |
| entity | tag |
| feature | column in a dynamic table or in a view |
Checkpoint 1 of 7· Match them up
Match each Feature Store object to the Snowflake object that implements it
Tap a term, then the definition that fits it.
Every Feature Store object is a native Snowflake object. That is why access control applies to them, and why dropping the schema deletes the whole feature store.
“A feature store in Snowflake is simply a schema.”Source: docs.snowflake.com
Sources1
2.Managed versus external feature views
You create a feature view in Python with snowflake.ml.feature_store.FeatureView. The constructor takes a Snowpark DataFrame that holds the feature logic, and that DataFrame must include the join-key columns of the entities you attach. If the features are time series, you also give a timestamp column.
A managed feature view stores its feature table in a dynamic table, and Snowflake refreshes it on your schedule, incrementally where it can. refresh_freq can be a time delta with a minimum of 1 minute, or a cron expression with a time zone. Incremental refresh needs change tracking on every source table. Snowflake tries to turn it on for you, which requires OWNERSHIP of the table. If you don't own the table, ask the owner to enable it or use refresh_mode='FULL'. Queries that a dynamic table can't maintain incrementally fall back to full refreshes, which add lag and cost.
Checkpoint 2 of 7· Fill the gap
Which parameter makes this feature view Snowflake-managed?
managed_fv = FeatureView(
name="MY_MANAGED_FV",
entities=[entity],
feature_df=my_df, # Snowpark DataFrame containing feature transformations
timestamp_col="ts", # optional timestamp column name in the dataframe
? ="5 minutes", # how often feature data refreshes
desc="my managed feature view" # optional description
)Setting refresh_freq designates the feature view as Snowflake-managed. refresh_mode='FULL' only controls how a refresh reads its sources.
Source: docs.snowflake.com| Aspect | Snowflake-managed | External |
|---|---|---|
| refresh_freq | Time delta or cron expression | None |
| Feature table | Dynamic table refreshed by Snowflake | Your own table; the feature view is a view on it |
| Who maintains features | Snowflake, on the schedule | You, for example with dbt |
| Storage | Dynamic table | No additional storage cost |
Checkpoint 3 of 7· Exam question
A data scientist is preparing a customer segmentation workload in Snowpark and a regularized logistic regression churn model on the same table, with columns such as `ANNUAL_INCOME` (tens of thousands) and `TENURE_MONTHS` (0 to 120). Select TWO of the following algorithms that require the numeric features to be scaled to produce reliable results.(Select 2)
Correct answers: B, C — K-means clustering, because assignments use Euclidean distance and the large-range income column would dominate the distance calculation.; L2-regularized logistic regression, because the penalty shrinks all coefficients equally and penalizes features measured in small units unfairly.
- A. Tree splits are threshold comparisons that are invariant to monotonic rescaling, so a random forest finds the same partitions whether or not columns are scaled.
- B. K-means groups rows by Euclidean distance, so an unscaled column with a large range effectively decides cluster membership on its own. Scaling puts every feature on a comparable footing.
- C. Regularization penalties act on coefficient size, which depends on feature units. Without scaling, features on small scales need large coefficients and get over-penalized.
- D. Information gain or Gini impurity is computed on class label distributions in each split, not on feature variance, so scaling has no effect on a single decision tree.
- E. Boosted trees also split on thresholds, so rescaling a column does not change the structure learned. Leaf weights depend on the target gradient rather than feature scale.
Sources2
3.Registering, versioning and updating feature views
Defining a feature view does not publish it. You publish it with fs.register_feature_view, giving it a version. block=True waits until the initial data is available, and overwrite=False stops you from replacing an existing feature view with the same name and version. Once registered, it can be retrieved with get_feature_view(name, version), listed with list_feature_views (filtered by entity or by name), or found in the Snowsight Feature Store UI and Universal Search. Calling attach_feature_desc before registering adds per-feature descriptions, which make features easier to find.
registered_fv: FeatureView = fs.register_feature_view(
feature_view=managed_fv, # feature view created above, could also use external_fv
version="1",
block=True, # whether function call blocks until initial data is available
overwrite=False, # whether to replace existing feature view with same name/version
)A registered pipeline definition is immutable, and that guarantees every consumer the same feature computation. update_feature_view can change only three properties: the refresh frequency, the warehouse that runs the transforms, and the description. To change a window length or a column, you register a new version and leave the old one running for the models that still read it.
Checkpoint 4 of 7· Check yourself
Version '1' of a feature view computes a 30-day average spend and feeds production models. You need a 60-day variant. What should you do?
Feature definitions cannot be modified after registration. update_feature_view only changes the refresh frequency, warehouse and description, so a new version is how you change features without disturbing existing consumers.
“Feature definitions and columns cannot be modified. To change the features in a feature store, create a new version of the feature view.”Source: docs.snowflake.com
Checkpoint 5 of 7· Exam question
A retail analytics team stores each customer's yearly spend across 40 product categories as one row of 40 numeric columns and wants to cluster customers by the shape of their spending mix, ignoring how large their total spend is. Which `snowflake.ml.modeling.preprocessing` transformer is the MOST appropriate preparation step?
Correct answer: A — `Normalizer` with `norm="l2"`, which rescales each row independently so every customer's spend vector has unit length.
- A. Normalizer works per sample rather than per column, so each customer's category vector gets unit norm and only the mix pattern remains, not the total magnitude.
- B. RobustScaler is column-wise and addresses outliers, not the row-level magnitude problem described, so large total spend still differentiates customers.
- C. MinMaxScaler also works per column. It aligns column ranges but leaves each customer's row length proportional to their overall spending level.
- D. StandardScaler is column-wise, so heavy spenders still have larger values across all categories in their row and total spend remains in the vectors.
Sources2
4.Point-in-time correct training datasets
To train a model you join features to a spine: a table of entity keys, label timestamps and labels. A naive join takes the latest feature value, which can come from after the label was observed. That leaks future information into training. The Feature Store avoids this by joining on time. Pass generate_dataset the spine together with spine_timestamp_col. For each spine row, the Feature Store runs an ASOF join and takes the most recent feature row for that entity at or before the spine timestamp.
The join needs a time column on each side. On the spine, that is the column you name in spine_timestamp_col. On the feature view, it is the timestamp_col you set when you created it, which is required for time-series features. Once training data comes from the Feature Store, ML Lineage records the path from source table to feature, dataset and trained model. Further details of the dataset object that generate_dataset returns are not covered in the documentation provided for this lesson.
Checkpoint 6 of 7· Check yourself
A spine has CUSTOMER_ID, LABEL_TS and CHURNED. The AVG_SPEND_30D feature view has a timestamp column. What does passing spine_timestamp_col='LABEL_TS' to generate_dataset achieve?
The ASOF join returns features as they were when each label was observed, which prevents future data from leaking into training.
“the Feature Store performs an ASOF join and selects the most recent snapshot row for each entity key at or before the spine timestamp”Source: docs.snowflake.com
Checkpoint 7 of 7· Exam question
A fraud team's `TRANSACTION_AMOUNT` column has a long tail: most values are under 200 but a few card-testing and bulk purchases exceed 500,000. The team wants a scaler for a distance-based model that is not dominated by those extremes. Which choice is MOST appropriate?
Correct answer: C — `RobustScaler`, which centers values on the median and divides by the interquartile range so extreme values barely influence the scale.
- A. Mean and standard deviation are strongly inflated by the extreme amounts, so ordinary transactions end up compressed around a similar slightly negative value.
- B. Dividing by the largest absolute value is determined entirely by the outlier, so typical amounts become close to zero and the problem remains.
- C. RobustScaler uses the median and IQR, statistics that are insensitive to a handful of extreme values, so typical transactions keep a meaningful spread.
- D. The maximum is set by the largest outlier, so nearly all ordinary transactions would be squeezed into a tiny band near zero.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Snowflake keeps an external feature view's features fresh, just without incremental refresh.Why is that wrong?
With refresh_freq=None the feature view is external, and you are responsible for maintaining the feature table, for example with dbt.
Covered in Managed versus external feature views
2.update_feature_view can change a registered feature view's transformation logic.Why is that wrong?
A registered pipeline definition is immutable. Only the refresh frequency, the warehouse and the description can be updated.
Covered in Registering, versioning and updating feature views
3.Joining the spine to the current feature table on the entity key gives valid training data.Why is that wrong?
Without the time-aware ASOF join that spine_timestamp_col enables, feature values from after the label time can leak into training.
Covered in Point-in-time correct training datasets
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“A feature view encapsulates a Python or SQL pipeline for transforming raw data into one or more related features.”
↩︎ What a feature store is in Snowflake“An entity is a higher-level abstraction that represents the subject matter of a feature.”
↩︎ What a feature store is in Snowflake“Backfill and point-in-time correct features with ASOF JOIN”
↩︎ What a feature store is in Snowflake“ML Lineage is automatically created when the feature store is used.”
↩︎ Point-in-time correct training datasets“A feature store in Snowflake is simply a schema.”
↩︎ Checkpoint - 2.
“The FeatureView constructor accepts a Snowpark DataFrame that contains the feature generation logic.”
↩︎ Managed versus external feature views“The value can be a time delta (minimum value 1 minute), or it can be a cron expression with time zone”
↩︎ Managed versus external feature views“To qualify for incremental refresh, each source table must have change tracking enabled.”
↩︎ Managed versus external feature views“External feature views are implemented as views on your feature table, so they incur no additional storage cost.”
↩︎ Managed versus external feature views“the table will be fully refreshed from the query at the specified frequency”
↩︎ Managed versus external feature views“Once a feature view has been completely defined, you can register it in the feature store using the feature store’s register_feature_view method”
↩︎ Registering, versioning and updating feature views“A timestamp column name is required if your feature view includes time-series features.”
↩︎ Point-in-time correct training datasets“If you don’t provide a schedule for refreshing the feature view, it’s considered external.”
↩︎ Exam trap 1“A feature view pipeline definition is immutable after it has been registered, providing consistent feature computation as long as the feature view exists.”
↩︎ Exam trap 2“A feature view is considered Snowflake-managed if you provide a schedule for refreshing it.”
↩︎ Prediction“Feature definitions and columns cannot be modified. To change the features in a feature store, create a new version of the feature view.”
↩︎ Checkpoint - 3.https://docs.snowflake.com/en/developer-guide/snowflake-ml/feature-store/advanced-feature-engineeringOfficial docs
“use generate_dataset with spine_timestamp_col to build training sets from the accumulated snapshots”
↩︎ Point-in-time correct training datasets“preventing future data from leaking into the model”
↩︎ Exam trap 3“the Feature Store performs an ASOF join and selects the most recent snapshot row for each entity key at or before the spine timestamp”
↩︎ Checkpoint