CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 1 · Lesson 8/48

    Train and Log a Feature Store Model with fe.log_model

    Train a model with features from a feature store table.

    9 min read
    2.08% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Train a model on the DataFrame returned by TrainingSet.load_df without breaking inference-time feature lookup
    • Log a model with FeatureEngineeringClient.log_model so that it carries its feature metadata
    • State the requirements and limits for feature store models, and where their lineage appears
    • Choose the correct client and package for Unity Catalog and for workspace feature tables

    1.Train on exactly what load_df returns

    A TrainingSet from fe.create_training_set is a specification. training_set.load_df() runs the joins and returns a Spark DataFrame with your label, any columns you kept, and one column per looked-up feature. You convert it to the format your framework expects, split it into inputs and label, and fit the model. The documentation's example puts the whole sequence inside with mlflow.start_run():, so the training and the logged model belong to the same MLflow run.

    Fitting a scikit-learn model on the loaded training set, then logging it with the Feature Engineering clientpython
      training_df = training_set.load_df().toPandas()
    
      # "training_df" columns ['total_purchases_30d', 'category', 'rating']
      X_train = training_df.drop(['rating'], axis=1)
      y_train = training_df.rating
    
      model = linear_model.LinearRegression().fit(X_train, y_train)
    
      fe.log_model(
        model=model,
        artifact_path="recommendation_model",
        flavor=mlflow.sklearn,
        training_set=training_set,
        registered_model_name="recommendation_model"
      )

    If you want the model to look up features automatically at inference, you must train on the DataFrame from load_df. When the model fetches features later, it rebuilds them from the recorded lookups and nothing else. So any change you make to that DataFrame before fitting is missing at inference, and the documentation warns that this decreases the model's performance. Separating the label column, as the example does with drop(['rating']), is how the inputs are formed. It does not change the feature values themselves.

    Checkpoint 1 of 5· Put it in order

    Put the steps for training a model with feature store features in order

    1. 1.Log the model with fe.log_model, passing training_set
    2. 2.Call fe.create_training_set with the label DataFrame, the lookups and the label column
    3. 3.Fit the model on that DataFrame
    4. 4.Call training_set.load_df() to get the joined DataFrame
    5. 5.Define a FeatureLookup for each feature table

    Sources1

    2.Log with the client, not plain MLflow

    The key call is fe.log_model, and its training_set= argument links the fitted model to the features it was trained on. The method comes from FeatureEngineeringClient, not from an MLflow flavor module. You still pass the MLflow flavor (flavor=mlflow.sklearn), and you can register the model in the same call with registered_model_name. When a model is logged this way, it keeps its feature references, and at inference time the model can optionally retrieve feature values automatically.

    That is the payoff of everything on the previous steps. A caller provides only the primary keys, such as customer_id and product_id, and the model fetches the rest. For batch inference the values come from the offline store and are joined to the new data before scoring. A model logged with a plain MLflow log_model call has no such link, so callers must assemble every feature themselves.

    Checkpoint 2 of 5· Fill the gap

    Which argument links the logged model to the features it was trained on?

      fe.log_model(
        model=model,
        artifact_path="recommendation_model",
        flavor=mlflow.sklearn,
         ? =training_set,
        registered_model_name="recommendation_model"
      )

    Checkpoint 3 of 5· Exam question

    A model was trained and logged using `fe.log_model` with its `TrainingSet`. A data scientist now needs to run batch inference on a new Spark DataFrame `inference_df` that only contains the primary key column `customer_id` plus any raw label columns, without manually re-joining the feature table. Which call correctly produces predictions with features re-attached automatically?

    Sources2

    3.Requirements, limits and lineage

    Inference-time feature lookup has three conditions, and the first two were covered above. You must log with the client's log_model. You must train on the unmodified output of load_df. Finally, the model type must have a corresponding python_flavor in MLflow. That covers scikit-learn, Keras, PyTorch, SparkML, LightGBM, XGBoost and TensorFlow Keras, as well as custom MLflow pyfunc models. Feature store models also work with the MLflow pyfunc interface, so MLflow can run batch inference with them.

    There is also a size limit: a model can use at most 50 tables and 100 functions for training. When you train and log through Feature Engineering in Unity Catalog, you also get lineage without extra work. Catalog Explorer shows which tables and functions the model was built from.

    Checkpoint 4 of 5· Check yourself

    After logging a model with fe.log_model on Unity Catalog feature tables, where can you see which feature tables were used to build it?

    Sources1

    4.FeatureEngineeringClient vs the legacy FeatureStoreClient

    Every call on this page also exists in a legacy form. FeatureStoreClient.create_training_set and FeatureStoreClient.log_model work against the Workspace Feature Store, which uses two-level table names such as recommender_system.customer_features. Which client you use depends on where the feature tables live. The package depends on your runtime version:

    Which package and Python client to use
    Databricks Runtime versionFeature tables inPackagePython client
    Databricks Runtime 14.3 ML and aboveUnity Catalogdatabricks-feature-engineeringFeatureEngineeringClient
    Databricks Runtime 14.3 ML and aboveWorkspacedatabricks-feature-engineeringFeatureStoreClient
    Databricks Runtime 14.2 ML and belowUnity Catalogdatabricks-feature-engineeringFeatureEngineeringClient
    Databricks Runtime 14.2 ML and belowWorkspacedatabricks-feature-storeFeatureStoreClient

    For Unity Catalog feature tables, the answer is always FeatureEngineeringClient from databricks-feature-engineering. The older databricks-feature-store package is deprecated. One version detail affects training runs: databricks-feature-engineering 0.7.0 and below does not work with MLflow 2.18.0 and above, so upgrade to 0.8.0 or later.

    Checkpoint 5 of 5· Check yourself

    Your feature tables are in Unity Catalog and you are on Databricks Runtime 15.4 ML. Which client should create the training set and log the model?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Preprocessing you apply to the load_df output is stored with the model and replayed at inference.Why is that wrong?

      Only the feature lookups are recorded. Changes you make to that DataFrame before training are not applied at inference, which decreases performance.

      Covered in Train on exactly what load_df returns

    2. 2.Logging a model with the MLflow flavor's own log_model, such as mlflow.sklearn.log_model, is enough for it to look up features at inference.Why is that wrong?

      Inference-time feature lookup requires logging through the Feature Engineering client's log_model, with the training set passed in.

      Covered in Log with the client, not plain MLflow

    Practise it for real

    Train and log a scikit-learn model whose features come from Unity Catalog feature tables

    1. 1.Create a FeatureLookup list that selects total_purchases_30d from ml.recommender_system.customer_features (key customer_id) and category from ml.recommender_system.product_features (key product_id).

      Why: Each lookup names a table, its features and the join key.

      You should see: A Python list of two FeatureLookup objects.

    2. 2.Inside with mlflow.start_run(), call fe.create_training_set(df=df, feature_lookups=feature_lookups, label='rating', exclude_columns=['customer_id', 'product_id']).

      Why: This declares the left join of the feature tables onto your labels and drops the ID columns from the output.

      You should see: A TrainingSet object. No data has been materialised yet.

    3. 3.Run training_set.load_df().toPandas(), then split it into X_train (all columns except rating) and y_train (rating), and fit LinearRegression.

      Why: The model must be trained on the DataFrame that load_df returns.

      You should see: A DataFrame with columns total_purchases_30d, category and rating, and a fitted model.

    4. 4.Call fe.log_model(model=model, artifact_path="recommendation_model", flavor=mlflow.sklearn, training_set=training_set, registered_model_name=...).

      Why: Passing training_set packages the feature references with the model.

      You should see: A logged and registered model that lists its source feature tables in its lineage in Catalog Explorer.

    Stuck? Get a nudge

    If create_training_set fails on a table with a DATE primary key, check whether that column was declared as a timeseries column.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “You must use the DataFrame returned by TrainingSet.load_df to train the model.”
      ↩︎ Train on exactly what load_df returns
      “The model type must have a corresponding python_flavor in MLflow.”
      ↩︎ Requirements, limits and lineage
      “A model can use at most 50 tables and 100 functions for training.”
      ↩︎ Requirements, limits and lineage
      “If you modify this DataFrame in any way before using it to train the model, the modifications are not applied”
      ↩︎ Exam trap 1
      “You must log the model using the log_model method of FeatureEngineeringClient”
      ↩︎ Exam trap 2
      “If you modify this DataFrame in any way before using it to train the model, the modifications are not applied”
      ↩︎ Prediction
      “Tables and functions that were used to create the model are automatically tracked and displayed.”
      ↩︎ Checkpoint
    2. 2.
      “the model retains references to these features. At inference time, the model can optionally retrieve feature values automatically.”
      ↩︎ Log with the client, not plain MLflow
      “In batch inference, feature values are retrieved from the offline store and joined with new data prior to scoring.”
      ↩︎ Log with the client, not plain MLflow
    3. 3.
      “databricks-feature-engineering<=0.7.0 is not compatible with mlflow>=2.18.0.”
      ↩︎ FeatureEngineeringClient vs the legacy FeatureStoreClient
      “As of version 0.17.0, databricks-feature-store has been deprecated.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.