CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 1 · Lesson 8/48

    Build a Training Set with FeatureLookup and create_training_set

    Train a model with features from a feature store table.

    13 min read
    2.08% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what a training set is and why a model trained on one keeps references to its features
    • Write FeatureLookup objects that select features from Unity Catalog feature tables
    • Predict how create_training_set joins feature tables to a label DataFrame, including key order and column naming
    • Use output_name, label=None, default_values and timestamp_lookup_key correctly

    Key concept

    Training set — A training set pairs your label DataFrame, which holds the raw rows, the labels and the lookup keys, with a list of features to pull from feature tables. Because the model is built from that specification and not from a hand-made join, it remembers which features it needs and where they come from.

    1.Why you don't just join the tables yourself

    In a Unity Catalog-enabled workspace, you can use any Delta table in Unity Catalog that has a primary key constraint as a feature table. You can read it with plain Spark SQL and join it to your labels, but then the model learns nothing about where its inputs came from. The Feature Engineering client gives you a different path. You first create a training dataset that declares which features to use and how to join them. When you train on that dataset, the model keeps references to those features.

    Those references matter later. When you score the model, it can fetch feature values from the feature store itself, so you only supply keys. Catalog Explorer also shows the model's lineage back to the tables it used. Everything in this lesson builds that declarative training dataset. You need two objects for it: one FeatureLookup per feature table, and one call to FeatureEngineeringClient.create_training_set.

    Checkpoint 1 of 7· Check yourself

    You define a training dataset with create_training_set and then train a model on it. What does the trained model keep?

    Sources12

    2.FeatureLookup: naming the features you want

    Many models can share one feature table, and few need every column. A FeatureLookup tells the client three things: the feature table's name, the features to take from it, and the key or keys used to join it to the DataFrame you pass to create_training_set. In Unity Catalog the table name has three levels (catalog.schema.table).

    Two FeatureLookups: two features from customer_features and one from product_featurespython
    feature_lookups = [
        FeatureLookup(
          table_name='ml.recommender_system.customer_features',
          feature_names=['total_purchases_30d', 'total_purchases_7d'],
          lookup_key='customer_id'
        ),
        FeatureLookup(
          table_name='ml.recommender_system.product_features',
          feature_names=['category'],
          lookup_key='product_id'
        )
      ]

    feature_names accepts a single name, a list of names, or None. None selects every feature except the primary keys, and the set is fixed when the training set is created. One training set can draw on several tables, but there is a ceiling: a model can use at most 50 tables and 100 functions for training.

    Checkpoint 2 of 7· Match them up

    Match each FeatureLookup argument to its role

    Tap a term, then the definition that fits it.

    Checkpoint 3 of 7· Exam question

    A team maintains a Unity Catalog feature table `customer_features` with primary key `customer_id`, containing columns `avg_purchase_value` and `days_since_last_order`. They have a separate DataFrame `labels_df` with `customer_id` and a `churned` label, and want to train a churn model using only the two named features joined by primary key while preserving feature lineage for the resulting model. Which approach best meets this requirement?

    Sources1

    3.create_training_set: keys, join type and excluded columns

    You pass the lookups to create_training_set together with the label DataFrame and the name of the label column. The DataFrame must contain a column for every primary key of every feature table you look up. The column names do not have to match the table's key names. The rule is about position: the type and order of the lookup_key columns must match the type and order of the table's primary keys, not counting timestamp keys. The documentation's example table has primary keys customer_id, dt, while the training DataFrame calls those columns cid and transaction_dt. Listing them in the same order is enough:

    A composite lookup_key whose column names differ from the table's primary keyspython
        FeatureLookup(
          table_name='ml.recommender_system.customer_features',
          feature_names=['total_purchases_30d', 'total_purchases_7d'],
          lookup_key=['cid', 'transaction_dt']
        ),

    Under the hood, the call performs a left join of each feature table onto your DataFrame. Every label row is kept. A row whose key has no match in a feature table gets an empty value for that table's features; it is not dropped. In the resulting DataFrame, all of your original columns stay, plus one column per looked-up feature. The exceptions are the columns you name in exclude_columns, which is typically used to drop ID columns that should not become model inputs:

    Building the training set and excluding the key columns from the outputpython
    training_set = fe.create_training_set(
      df=training_df,
      feature_lookups=feature_lookups,
      label='rating',
      exclude_columns=['customer_id', 'product_id']
    )

    Checkpoint 4 of 7· Check yourself

    Your label DataFrame has 10,000 rows. 300 of them have a customer_id that does not exist in the customer feature table. How many rows does create_training_set return?

    Sources21

    4.Renaming outputs, unsupervised sets and default values

    Three options handle the cases that a basic lookup cannot.

    **output_name** replaces the feature's name in the DataFrame returned by load_df. You need it when the same feature appears twice. For example, one temperature feature may be looked up once by pickup ZIP code and once by drop-off ZIP code. Each of those lookups needs its own output_name, otherwise both would produce a column with the same name:

    The same feature joined on two different keys, each with a unique output_namepython
        FeatureLookup(
          table_name='ml.taxi_data.zip_features',
          feature_names=['temperature'],
          lookup_key=['pickup_zip'],
          output_name='pickup_temp'
        ),
        FeatureLookup(
          table_name='ml.taxi_data.zip_features',
          feature_names=['temperature'],
          lookup_key=['dropoff_zip'],
          output_name='dropoff_temp'
        )

    **label=None** builds a training set with no label column. Use it for unsupervised models, such as clustering customers by their interests.

    **default_values** fills in a value when the feature store has no computed value for an ID. These are exactly the gaps the left join leaves. You pass it as a dictionary keyed by feature name. If you renamed the columns with rename_outputs, the dictionary must use the new names.

    Default values for features that may be missing for some customerspython
            lookup_key="customer_id",
            default_values={
              "age": 18,
              "membership_tier": "bronze"
            },

    Checkpoint 5 of 7· Check yourself

    You want both the pickup and the drop-off temperature from the same zip_features table. What is the correct approach?

    Checkpoint 6 of 7· Exam question

    After training a scikit-learn model on a training set created via `fe.create_training_set(...)`, a data scientist wants Model Serving to automatically look up the same features from the feature table at inference time, so client requests only need to supply `customer_id` rather than every feature value. Which step accomplishes this?

    Sources1

    5.Time series tables: timestamp_lookup_key, not lookup_key

    Some features change over time. If a label recorded at 8:50 is joined to a sensor reading taken at 8:52, information from the future leaks into training. Point-in-time correctness prevents this kind of data leakage. A table becomes a time series feature table when its timestamp primary key is declared as a timeseries column. Every FeatureLookup on such a table must be a point-in-time lookup.

    You set this up with a separate argument. lookup_key names the ordinary key columns. timestamp_lookup_key names the DataFrame column that holds each row's timestamp. For each row, the feature store returns the latest feature values from before that timestamp, or null if there are none. Never put the timestamp column in lookup_key.

    A point-in-time lookup: the user features are matched as of each ad impressionpython
      FeatureLookup(
        table_name="ml.ads_team.user_features",
        feature_names=["purchases_30d", "is_free_trial_active"],
        lookup_key="u_id",
        timestamp_lookup_key="ad_impression_ts"
      ),

    One error catches people out. If a feature table has a DATE or TIMESTAMP primary key that is *not* declared as a timeseries column, create_training_set returns an error. There are two fixes. Declare the column as a timeseries column if you want point-in-time behaviour. Change its type to STRING if it is only an exact-match key.

    Checkpoint 7 of 7· Check yourself

    A feature table has primary keys (user_id, event_date), where event_date is a DATE column that was not declared as a timeseries column. You call create_training_set on it. What happens?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The lookup_key columns must have the same names as the feature table's primary keys.Why is that wrong?

      Names do not matter. Type and order do. The lookup_key columns are matched by position against the primary keys (cid matches customer_id).

      Covered in create_training_set: keys, join type and excluded columns

    2. 2.To join a time series table, add the timestamp column to lookup_key.Why is that wrong?

      Timestamp columns go in timestamp_lookup_key. That is what triggers the as-of lookup of the latest earlier value.

      Covered in Time series tables: timestamp_lookup_key, not lookup_key

    3. 3.When features have been renamed, default_values still uses the original feature names.Why is that wrong?

      Once rename_outputs is used, default_values must refer to the renamed columns.

      Covered in Renaming outputs, unsupervised sets and default values

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “you first create a training dataset that defines the features to use and how to join them”
      ↩︎ Why you don't just join the tables yourself
      “A model can use at most 50 tables and 100 functions for training.”
      ↩︎ FeatureLookup: naming the features you want
      “It preserves all columns of the DataFrame provided to create_training_set except those excluded using exclude_columns.”
      ↩︎ create_training_set: keys, join type and excluded columns
      “Set label=None when creating a TrainingSet for unsupervised learning models.”
      ↩︎ Renaming outputs, unsupervised sets and default values
      “you can specify default values for features to handle cases where the Feature Store does not have a computed feature value for an ID.”
      ↩︎ Renaming outputs, unsupervised sets and default values
      “The type and order of lookup_key columns in the DataFrame must match the type and order of the primary keys”
      ↩︎ Exam trap 1
      “Timestamp columns should not be used as a lookup_key.”
      ↩︎ Exam trap 2
      “If the feature columns are renamed using the rename_outputs parameter, default_values must use the renamed feature names.”
      ↩︎ Exam trap 3
      “When you then train the model, it retains references to those features.”
      ↩︎ Checkpoint
      “feature_names takes a single feature name, a list of feature names, or None to look up all features (excluding primary keys)”
      ↩︎ Prediction
      “including the name of the feature table, the name(s) of the features, and the key(s) to use when joining the feature table”
      ↩︎ Checkpoint
      “it creates a training dataset by performing a left join”
      ↩︎ Checkpoint
      “Use a unique output_name for each FeatureLookup output.”
      ↩︎ Checkpoint
      “DATE or TIMESTAMP columns used as primary keys must be declared as timeseries keys.”
      ↩︎ Checkpoint
    2. 2.
      “you can use any Delta table in Unity Catalog that includes a primary key constraint as a feature table”
      ↩︎ Why you don't just join the tables yourself
      “The training data must contain column(s) corresponding to each of the primary keys of the feature tables.”
      ↩︎ create_training_set: keys, join type and excluded columns
      “A training set consists of a list of features and a DataFrame containing raw training data, labels, and primary keys”
      ↩︎ Key concept
    3. 3.
      “Any FeatureLookup on a time series feature table must be a point-in-time lookup, so it must specify a timestamp_lookup_key”
      ↩︎ Time series tables: timestamp_lookup_key, not lookup_key
      “Databricks Feature Store retrieves the latest feature values prior to the timestamps specified in the DataFrame's timestamp_lookup_key column”
      ↩︎ Time series tables: timestamp_lookup_key, not lookup_key
      “This is important to prevent data leakage”
      ↩︎ Time series tables: timestamp_lookup_key, not lookup_key

    Continue to page 2 of 2

    Train and Log a Feature Store Model with fe.log_model

    Spotted a mistake, or was something unclear? Tell us.