What you will be able to do
- Explain what a training set is and why a model trained on one keeps references to its features
- Write FeatureLookup objects that select features from Unity Catalog feature tables
- Predict how create_training_set joins feature tables to a label DataFrame, including key order and column naming
- Use output_name, label=None, default_values and timestamp_lookup_key correctly
Key concept
Training set — A training set pairs your label DataFrame, which holds the raw rows, the labels and the lookup keys, with a list of features to pull from feature tables. Because the model is built from that specification and not from a hand-made join, it remembers which features it needs and where they come from.
1.Why you don't just join the tables yourself
In a Unity Catalog-enabled workspace, you can use any Delta table in Unity Catalog that has a primary key constraint as a feature table. You can read it with plain Spark SQL and join it to your labels, but then the model learns nothing about where its inputs came from. The Feature Engineering client gives you a different path. You first create a training dataset that declares which features to use and how to join them. When you train on that dataset, the model keeps references to those features.
Those references matter later. When you score the model, it can fetch feature values from the feature store itself, so you only supply keys. Catalog Explorer also shows the model's lineage back to the tables it used. Everything in this lesson builds that declarative training dataset. You need two objects for it: one FeatureLookup per feature table, and one call to FeatureEngineeringClient.create_training_set.
It holds no references to the features it used. So it cannot look up feature values from the feature store by itself at inference time. Every caller has to rebuild the same join, which the feature store exists to avoid.
Checkpoint 1 of 7· Check yourself
You define a training dataset with create_training_set and then train a model on it. What does the trained model keep?
Training on a dataset that declares its features means the model retains references to those features, which is what later allows feature lookup at inference.
“When you then train the model, it retains references to those features.”Source: docs.databricks.com
2.FeatureLookup: naming the features you want
Many models can share one feature table, and few need every column. A FeatureLookup tells the client three things: the feature table's name, the features to take from it, and the key or keys used to join it to the DataFrame you pass to create_training_set. In Unity Catalog the table name has three levels (catalog.schema.table).
feature_lookups = [
FeatureLookup(
table_name='ml.recommender_system.customer_features',
feature_names=['total_purchases_30d', 'total_purchases_7d'],
lookup_key='customer_id'
),
FeatureLookup(
table_name='ml.recommender_system.product_features',
feature_names=['category'],
lookup_key='product_id'
)
]feature_names accepts a single name, a list of names, or None. None selects every feature except the primary keys, and the set is fixed when the training set is created. One training set can draw on several tables, but there is a ceiling: a model can use at most 50 tables and 100 functions for training.
Checkpoint 2 of 7· Match them up
Match each FeatureLookup argument to its role
Tap a term, then the definition that fits it.
A FeatureLookup names the table, the features and the join keys. Nothing else is required to select features.
“including the name of the feature table, the name(s) of the features, and the key(s) to use when joining the feature table”Source: docs.databricks.com
Checkpoint 3 of 7· Exam question
A team maintains a Unity Catalog feature table `customer_features` with primary key `customer_id`, containing columns `avg_purchase_value` and `days_since_last_order`. They have a separate DataFrame `labels_df` with `customer_id` and a `churned` label, and want to train a churn model using only the two named features joined by primary key while preserving feature lineage for the resulting model. Which approach best meets this requirement?
Correct answer: A — Instantiate a `FeatureEngineeringClient`, call `create_training_set` on `labels_df` with a `FeatureLookup` for `customer_features` naming `customer_id` as the lookup key and the two feature columns, then call `.load_df()` to materialize the joined training DataFrame.
- A. Using `create_training_set` with a `FeatureLookup` automatically joins the named features by key and, critically, embeds lineage metadata so the resulting model tracks which feature table and columns it depends on.
- B. Manually joining `read_table` output like this returns the correct feature values but bypasses `FeatureLookup`, so the resulting model has no recorded lineage back to the feature table for later governance.
- C. A `FeatureFunction` is for deriving new features computed at request time from input arguments, not for pulling existing stored columns, so it doesn't fit a scenario that just needs two already-stored features by key.
- D. Round-tripping the feature table through a CSV export defeats the purpose of a governed feature store and loses the primary-key-based lookup and lineage tracking that Feature Engineering in Unity Catalog provides.
Sources1
3.create_training_set: keys, join type and excluded columns
You pass the lookups to create_training_set together with the label DataFrame and the name of the label column. The DataFrame must contain a column for every primary key of every feature table you look up. The column names do not have to match the table's key names. The rule is about position: the type and order of the lookup_key columns must match the type and order of the table's primary keys, not counting timestamp keys. The documentation's example table has primary keys customer_id, dt, while the training DataFrame calls those columns cid and transaction_dt. Listing them in the same order is enough:
FeatureLookup(
table_name='ml.recommender_system.customer_features',
feature_names=['total_purchases_30d', 'total_purchases_7d'],
lookup_key=['cid', 'transaction_dt']
),Under the hood, the call performs a left join of each feature table onto your DataFrame. Every label row is kept. A row whose key has no match in a feature table gets an empty value for that table's features; it is not dropped. In the resulting DataFrame, all of your original columns stay, plus one column per looked-up feature. The exceptions are the columns you name in exclude_columns, which is typically used to drop ID columns that should not become model inputs:
training_set = fe.create_training_set(
df=training_df,
feature_lookups=feature_lookups,
label='rating',
exclude_columns=['customer_id', 'product_id']
)Checkpoint 4 of 7· Check yourself
Your label DataFrame has 10,000 rows. 300 of them have a customer_id that does not exist in the customer feature table. How many rows does create_training_set return?
create_training_set left-joins each feature table onto the label DataFrame. Every label row is kept, and unmatched rows have no value for that table's features.
“it creates a training dataset by performing a left join”Source: docs.databricks.com
4.Renaming outputs, unsupervised sets and default values
Three options handle the cases that a basic lookup cannot.
**output_name** replaces the feature's name in the DataFrame returned by load_df. You need it when the same feature appears twice. For example, one temperature feature may be looked up once by pickup ZIP code and once by drop-off ZIP code. Each of those lookups needs its own output_name, otherwise both would produce a column with the same name:
FeatureLookup(
table_name='ml.taxi_data.zip_features',
feature_names=['temperature'],
lookup_key=['pickup_zip'],
output_name='pickup_temp'
),
FeatureLookup(
table_name='ml.taxi_data.zip_features',
feature_names=['temperature'],
lookup_key=['dropoff_zip'],
output_name='dropoff_temp'
)**label=None** builds a training set with no label column. Use it for unsupervised models, such as clustering customers by their interests.
**default_values** fills in a value when the feature store has no computed value for an ID. These are exactly the gaps the left join leaves. You pass it as a dictionary keyed by feature name. If you renamed the columns with rename_outputs, the dictionary must use the new names.
lookup_key="customer_id",
default_values={
"age": 18,
"membership_tier": "bronze"
},Checkpoint 5 of 7· Check yourself
You want both the pickup and the drop-off temperature from the same zip_features table. What is the correct approach?
A list of lookup keys means a composite key matched against the primary keys in order. It does not mean two separate joins. Two lookups with distinct output names give two separate columns.
“Use a unique output_name for each FeatureLookup output.”Source: docs.databricks.com
Checkpoint 6 of 7· Exam question
After training a scikit-learn model on a training set created via `fe.create_training_set(...)`, a data scientist wants Model Serving to automatically look up the same features from the feature table at inference time, so client requests only need to supply `customer_id` rather than every feature value. Which step accomplishes this?
Correct answer: A — Pass the original `TrainingSet` object returned by `create_training_set` to `fe.log_model(model=model, artifact_path='model', flavor=mlflow.sklearn, training_set=training_set, registered_model_name=...)`, so the logged model packages the feature lookup metadata needed for automatic retrieval.
- A. Passing the `TrainingSet` into `fe.log_model` packages the `FeatureLookup` definitions with the model, so a compatible serving environment can fetch the current feature values by key automatically rather than requiring the caller to supply them.
- B. Plain `mlflow.sklearn.log_model` has no awareness of feature lookups, so the resulting model still expects every feature value in the request payload; documentation elsewhere does not change what the endpoint accepts.
- C. Storing the training DataFrame as a static artifact captures only a snapshot of features at training time and provides no mechanism for a serving endpoint to perform fresh, key-based lookups on each request.
- D. Registering directly from the artifact URI still yields a model expecting full feature vectors as input, and manually wiring `fe.read_table` calls into the endpoint is not how Feature Engineering automatic lookup is configured.
Sources1
5.Time series tables: timestamp_lookup_key, not lookup_key
Some features change over time. If a label recorded at 8:50 is joined to a sensor reading taken at 8:52, information from the future leaks into training. Point-in-time correctness prevents this kind of data leakage. A table becomes a time series feature table when its timestamp primary key is declared as a timeseries column. Every FeatureLookup on such a table must be a point-in-time lookup.
You set this up with a separate argument. lookup_key names the ordinary key columns. timestamp_lookup_key names the DataFrame column that holds each row's timestamp. For each row, the feature store returns the latest feature values from before that timestamp, or null if there are none. Never put the timestamp column in lookup_key.
FeatureLookup(
table_name="ml.ads_team.user_features",
feature_names=["purchases_30d", "is_free_trial_active"],
lookup_key="u_id",
timestamp_lookup_key="ad_impression_ts"
),One error catches people out. If a feature table has a DATE or TIMESTAMP primary key that is *not* declared as a timeseries column, create_training_set returns an error. There are two fixes. Declare the column as a timeseries column if you want point-in-time behaviour. Change its type to STRING if it is only an exact-match key.
Checkpoint 7 of 7· Check yourself
A feature table has primary keys (user_id, event_date), where event_date is a DATE column that was not declared as a timeseries column. You call create_training_set on it. What happens?
DATE and TIMESTAMP primary keys must be declared as timeseries keys. Otherwise create_training_set, create_feature_spec and publish_table fail.
“DATE or TIMESTAMP columns used as primary keys must be declared as timeseries keys.”Source: docs.databricks.com
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The lookup_key columns must have the same names as the feature table's primary keys.Why is that wrong?
Names do not matter. Type and order do. The lookup_key columns are matched by position against the primary keys (cid matches customer_id).
Covered in create_training_set: keys, join type and excluded columns
2.To join a time series table, add the timestamp column to lookup_key.Why is that wrong?
Timestamp columns go in timestamp_lookup_key. That is what triggers the as-of lookup of the latest earlier value.
Covered in Time series tables: timestamp_lookup_key, not lookup_key
3.When features have been renamed, default_values still uses the original feature names.Why is that wrong?
Once rename_outputs is used, default_values must refer to the renamed columns.
Covered in Renaming outputs, unsupervised sets and default values
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/feature-store/train-models-with-feature-storeOfficial docs
“you first create a training dataset that defines the features to use and how to join them”
↩︎ Why you don't just join the tables yourself“A model can use at most 50 tables and 100 functions for training.”
↩︎ FeatureLookup: naming the features you want“It preserves all columns of the DataFrame provided to create_training_set except those excluded using exclude_columns.”
↩︎ create_training_set: keys, join type and excluded columns“Set label=None when creating a TrainingSet for unsupervised learning models.”
↩︎ Renaming outputs, unsupervised sets and default values“you can specify default values for features to handle cases where the Feature Store does not have a computed feature value for an ID.”
↩︎ Renaming outputs, unsupervised sets and default values“The type and order of lookup_key columns in the DataFrame must match the type and order of the primary keys”
↩︎ Exam trap 1“Timestamp columns should not be used as a lookup_key.”
↩︎ Exam trap 2“If the feature columns are renamed using the rename_outputs parameter, default_values must use the renamed feature names.”
↩︎ Exam trap 3“When you then train the model, it retains references to those features.”
↩︎ Checkpoint“feature_names takes a single feature name, a list of feature names, or None to look up all features (excluding primary keys)”
↩︎ Prediction“including the name of the feature table, the name(s) of the features, and the key(s) to use when joining the feature table”
↩︎ Checkpoint“it creates a training dataset by performing a left join”
↩︎ Checkpoint“Use a unique output_name for each FeatureLookup output.”
↩︎ Checkpoint“DATE or TIMESTAMP columns used as primary keys must be declared as timeseries keys.”
↩︎ Checkpoint - 2.
“you can use any Delta table in Unity Catalog that includes a primary key constraint as a feature table”
↩︎ Why you don't just join the tables yourself“The training data must contain column(s) corresponding to each of the primary keys of the feature tables.”
↩︎ create_training_set: keys, join type and excluded columns“A training set consists of a list of features and a DataFrame containing raw training data, labels, and primary keys”
↩︎ Key concept - 3.
“Any FeatureLookup on a time series feature table must be a point-in-time lookup, so it must specify a timestamp_lookup_key”
↩︎ Time series tables: timestamp_lookup_key, not lookup_key“Databricks Feature Store retrieves the latest feature values prior to the timestamps specified in the DataFrame's timestamp_lookup_key column”
↩︎ Time series tables: timestamp_lookup_key, not lookup_key“This is important to prevent data leakage”
↩︎ Time series tables: timestamp_lookup_key, not lookup_key