What you will be able to do
- Train a model on the DataFrame returned by TrainingSet.load_df without breaking inference-time feature lookup
- Log a model with FeatureEngineeringClient.log_model so that it carries its feature metadata
- State the requirements and limits for feature store models, and where their lineage appears
- Choose the correct client and package for Unity Catalog and for workspace feature tables
1.Train on exactly what load_df returns
A TrainingSet from fe.create_training_set is a specification. training_set.load_df() runs the joins and returns a Spark DataFrame with your label, any columns you kept, and one column per looked-up feature. You convert it to the format your framework expects, split it into inputs and label, and fit the model. The documentation's example puts the whole sequence inside with mlflow.start_run():, so the training and the logged model belong to the same MLflow run.
training_df = training_set.load_df().toPandas()
# "training_df" columns ['total_purchases_30d', 'category', 'rating']
X_train = training_df.drop(['rating'], axis=1)
y_train = training_df.rating
model = linear_model.LinearRegression().fit(X_train, y_train)
fe.log_model(
model=model,
artifact_path="recommendation_model",
flavor=mlflow.sklearn,
training_set=training_set,
registered_model_name="recommendation_model"
)If you want the model to look up features automatically at inference, you must train on the DataFrame from load_df. When the model fetches features later, it rebuilds them from the recorded lookups and nothing else. So any change you make to that DataFrame before fitting is missing at inference, and the documentation warns that this decreases the model's performance. Separating the label column, as the example does with drop(['rating']), is how the inputs are formed. It does not change the feature values themselves.
Checkpoint 1 of 5· Put it in order
Put the steps for training a model with feature store features in order
- 1.Log the model with fe.log_model, passing training_set
- 2.Call fe.create_training_set with the label DataFrame, the lookups and the label column
- 3.Fit the model on that DataFrame
- 4.Call training_set.load_df() to get the joined DataFrame
- 5.Define a FeatureLookup for each feature table
The lookups feed the training set. load_df produces the data the model must be trained on, and log_model records the training set with the fitted model.
“You must use the DataFrame returned by TrainingSet.load_df to train the model.”Source: docs.databricks.com
Sources1
2.Log with the client, not plain MLflow
The key call is fe.log_model, and its training_set= argument links the fitted model to the features it was trained on. The method comes from FeatureEngineeringClient, not from an MLflow flavor module. You still pass the MLflow flavor (flavor=mlflow.sklearn), and you can register the model in the same call with registered_model_name. When a model is logged this way, it keeps its feature references, and at inference time the model can optionally retrieve feature values automatically.
That is the payoff of everything on the previous steps. A caller provides only the primary keys, such as customer_id and product_id, and the model fetches the rest. For batch inference the values come from the offline store and are joined to the new data before scoring. A model logged with a plain MLflow log_model call has no such link, so callers must assemble every feature themselves.
Checkpoint 2 of 5· Fill the gap
Which argument links the logged model to the features it was trained on?
fe.log_model(
model=model,
artifact_path="recommendation_model",
flavor=mlflow.sklearn,
? =training_set,
registered_model_name="recommendation_model"
)Passing the TrainingSet object as training_set is what packages the feature metadata with the model. The loaded DataFrame, or the lookups on their own, would not do this.
Source: docs.databricks.comCheckpoint 3 of 5· Exam question
A model was trained and logged using `fe.log_model` with its `TrainingSet`. A data scientist now needs to run batch inference on a new Spark DataFrame `inference_df` that only contains the primary key column `customer_id` plus any raw label columns, without manually re-joining the feature table. Which call correctly produces predictions with features re-attached automatically?
Correct answer: A — `fe.score_batch(model_uri='models:/churn_model/3', df=inference_df)`, which reads the feature lookup metadata stored with the model to join the current feature values by `customer_id` before scoring.
- A. `score_batch` reads the `FeatureLookup` metadata embedded in the model by `fe.log_model` and performs the join against the current feature table automatically, which is exactly what is needed here.
- B. A plain pyfunc model has no built-in awareness of feature tables or lookup keys, so calling `predict` on a DataFrame missing feature columns raises a schema mismatch rather than triggering an automatic join.
- C. Manually reading the feature table and joining it works mechanically, but it discards the point of `fe.log_model` and `score_batch`, which is to avoid hand-writing the same join logic at every inference call site.
- D. Selecting the full feature table and filtering it without an explicit join key still requires manual alignment with `inference_df`, and passing it straight into a plain sklearn model without a join reintroduces exactly the manual work the feature store workflow avoids.
Sources2
3.Requirements, limits and lineage
Inference-time feature lookup has three conditions, and the first two were covered above. You must log with the client's log_model. You must train on the unmodified output of load_df. Finally, the model type must have a corresponding python_flavor in MLflow. That covers scikit-learn, Keras, PyTorch, SparkML, LightGBM, XGBoost and TensorFlow Keras, as well as custom MLflow pyfunc models. Feature store models also work with the MLflow pyfunc interface, so MLflow can run batch inference with them.
There is also a size limit: a model can use at most 50 tables and 100 functions for training. When you train and log through Feature Engineering in Unity Catalog, you also get lineage without extra work. Catalog Explorer shows which tables and functions the model was built from.
Checkpoint 4 of 5· Check yourself
After logging a model with fe.log_model on Unity Catalog feature tables, where can you see which feature tables were used to build it?
Feature Engineering in Unity Catalog tracks the model's lineage automatically and shows it in Catalog Explorer.
“Tables and functions that were used to create the model are automatically tracked and displayed.”Source: docs.databricks.com
Sources1
4.FeatureEngineeringClient vs the legacy FeatureStoreClient
Every call on this page also exists in a legacy form. FeatureStoreClient.create_training_set and FeatureStoreClient.log_model work against the Workspace Feature Store, which uses two-level table names such as recommender_system.customer_features. Which client you use depends on where the feature tables live. The package depends on your runtime version:
| Databricks Runtime version | Feature tables in | Package | Python client |
|---|---|---|---|
| Databricks Runtime 14.3 ML and above | Unity Catalog | databricks-feature-engineering | FeatureEngineeringClient |
| Databricks Runtime 14.3 ML and above | Workspace | databricks-feature-engineering | FeatureStoreClient |
| Databricks Runtime 14.2 ML and below | Unity Catalog | databricks-feature-engineering | FeatureEngineeringClient |
| Databricks Runtime 14.2 ML and below | Workspace | databricks-feature-store | FeatureStoreClient |
For Unity Catalog feature tables, the answer is always FeatureEngineeringClient from databricks-feature-engineering. The older databricks-feature-store package is deprecated. One version detail affects training runs: databricks-feature-engineering 0.7.0 and below does not work with MLflow 2.18.0 and above, so upgrade to 0.8.0 or later.
Checkpoint 5 of 5· Check yourself
Your feature tables are in Unity Catalog and you are on Databricks Runtime 15.4 ML. Which client should create the training set and log the model?
Unity Catalog feature tables always use FeatureEngineeringClient. FeatureStoreClient is only for the legacy Workspace Feature Store, and databricks-feature-store is deprecated.
“As of version 0.17.0, databricks-feature-store has been deprecated.”Source: docs.databricks.com
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Preprocessing you apply to the load_df output is stored with the model and replayed at inference.Why is that wrong?
Only the feature lookups are recorded. Changes you make to that DataFrame before training are not applied at inference, which decreases performance.
Covered in Train on exactly what load_df returns
2.Logging a model with the MLflow flavor's own log_model, such as mlflow.sklearn.log_model, is enough for it to look up features at inference.Why is that wrong?
Inference-time feature lookup requires logging through the Feature Engineering client's log_model, with the training set passed in.
Covered in Log with the client, not plain MLflow
Practise it for real
Train and log a scikit-learn model whose features come from Unity Catalog feature tables
1.Create a FeatureLookup list that selects total_purchases_30d from ml.recommender_system.customer_features (key customer_id) and category from ml.recommender_system.product_features (key product_id).
Why: Each lookup names a table, its features and the join key.
You should see: A Python list of two FeatureLookup objects.
2.Inside with mlflow.start_run(), call fe.create_training_set(df=df, feature_lookups=feature_lookups, label='rating', exclude_columns=['customer_id', 'product_id']).
Why: This declares the left join of the feature tables onto your labels and drops the ID columns from the output.
You should see: A TrainingSet object. No data has been materialised yet.
3.Run training_set.load_df().toPandas(), then split it into X_train (all columns except rating) and y_train (rating), and fit LinearRegression.
Why: The model must be trained on the DataFrame that load_df returns.
You should see: A DataFrame with columns total_purchases_30d, category and rating, and a fitted model.
4.Call fe.log_model(model=model, artifact_path="recommendation_model", flavor=mlflow.sklearn, training_set=training_set, registered_model_name=...).
Why: Passing training_set packages the feature references with the model.
You should see: A logged and registered model that lists its source feature tables in its lineage in Catalog Explorer.
Stuck? Get a nudge
If create_training_set fails on a table with a DATE primary key, check whether that column was declared as a timeseries column.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/feature-store/train-models-with-feature-storeOfficial docs
“You must use the DataFrame returned by TrainingSet.load_df to train the model.”
↩︎ Train on exactly what load_df returns“The model type must have a corresponding python_flavor in MLflow.”
↩︎ Requirements, limits and lineage“A model can use at most 50 tables and 100 functions for training.”
↩︎ Requirements, limits and lineage“If you modify this DataFrame in any way before using it to train the model, the modifications are not applied”
↩︎ Exam trap 1“You must log the model using the log_model method of FeatureEngineeringClient”
↩︎ Exam trap 2“If you modify this DataFrame in any way before using it to train the model, the modifications are not applied”
↩︎ Prediction“Tables and functions that were used to create the model are automatically tracked and displayed.”
↩︎ Checkpoint - 2.
“the model retains references to these features. At inference time, the model can optionally retrieve feature values automatically.”
↩︎ Log with the client, not plain MLflow“In batch inference, feature values are retrieved from the offline store and joined with new data prior to scoring.”
↩︎ Log with the client, not plain MLflow - 3.
“databricks-feature-engineering<=0.7.0 is not compatible with mlflow>=2.18.0.”
↩︎ FeatureEngineeringClient vs the legacy FeatureStoreClient“As of version 0.17.0, databricks-feature-store has been deprecated.”
↩︎ Checkpoint