CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 1 · Lesson 9/48

    Batch Scoring with fe.score_batch and Feature Store Tables

    Score a model using features from a feature store table.

    10 min read
    2.08% of exam
    2 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why a model logged with FeatureEngineeringClient.log_model needs only primary keys at scoring time
    • Call fe.score_batch on a DataFrame of keys and predict the columns of the returned DataFrame
    • Override a packaged feature with a custom value and supply columns that came from outside the feature store
    • Predict what score_batch returns when a lookup key does not exist in the feature table

    Key concept

    Model packaged with feature metadata — A model logged through the Feature Engineering client together with its training set keeps a record of which feature-table features it was trained on. When you score it, you pass only the entity keys, and the model fetches and joins the feature values itself.

    1.Why scoring needs only primary keys

    Scoring a model with feature store features depends on one decision made earlier, when the model was logged. If you log a model with FeatureEngineeringClient.log_model and pass the training_set used to build its training data, the model is packaged with feature metadata. It records which tables and features it was trained on. The training set itself belongs to a separate lesson. What matters here is the result: at inference time the model can look those features up again by itself.

    Logging the model with its training set is what packages the feature metadata (Feature Engineering in Unity Catalog)python
      fe.log_model(
        model=model,
        artifact_path="recommendation_model",
        flavor=mlflow.sklearn,
        training_set=training_set,
        registered_model_name="recommendation_model"
      )

    That packaging changes what the caller has to send. The caller doesn't rebuild the feature joins in the scoring pipeline. It sends the entity keys, such as customer_id, and the model fetches the rest. For batch scoring, the documentation says where those values come from: they are read from the offline store, which holds feature tables as Delta tables, and joined to the new data before the model makes predictions. Real-time serving reads from an online store instead, but this page is about batch scoring. The same computed feature values serve both training and scoring, so there's no second copy of the feature logic in the inference code that could drift.

    Checkpoint 1 of 6· Check yourself

    A model was trained on features from ml.recommender_system.customer_features and logged with fe.log_model(..., training_set=training_set). In a batch scoring job, where do the feature values come from?

    Sources1

    2.Calling fe.score_batch and reading its output

    The batch scoring call is FeatureEngineeringClient.score_batch. It takes two arguments. model_uri points to the logged or registered model, and df is the DataFrame to score. If the model is packaged with features, the call retrieves the required features before it scores the model. The returned DataFrame doesn't hold only predictions. It is your input with extra columns added: one for each looked-up feature value, plus a prediction column.

    Batch inference with automatic feature lookup, with the returned columns noted in the commentspython
    fe = FeatureEngineeringClient()
    
    # batch_df has columns 'customer_id' and 'product_id'
    predictions = fe.score_batch(
        model_uri=model_uri,
        df=batch_df
    )
    
    # The 'predictions' DataFrame has these columns:
    # 'customer_id', 'product_id', 'total_purchases_30d', 'category', 'prediction'

    The model_uri is usually a registry URI. The documentation's examples use the form models:/<name>/<version>, for example models:/ban_prediction_model/1. Because the output keeps the keys and the feature values next to each prediction, you can check exactly which feature values led to each score without a second join.

    Checkpoint 2 of 6· Fill the gap

    Which FeatureEngineeringClient method completes this batch inference call?

    # batch_df has columns ['customer_id', 'account_creation_date']
    predictions = fe. ? (
      model_uri='models:/ban_prediction_model/1',
      df=batch_df
    )

    Checkpoint 3 of 6· Exam question

    A ML engineer at a retail company trained and logged a model using `FeatureEngineeringClient.log_model()`, embedding feature lookup metadata from a Unity Catalog feature table. Now they need to generate predictions for a new batch of records that lack the derived feature columns used during training. Which approach lets them generate predictions without manually joining the feature table into the batch DataFrame?

    Sources2

    3.Overriding features and supplying non-feature-store columns

    Automatic lookup is the default, but you can turn it off for individual features. To score with your own value for a feature, add a column with that feature's name to the DataFrame you pass to score_batch. The API uses your values for that feature and looks up only the features you didn't supply. This is useful for what-if scoring or for correcting a value you know is out of date in the table.

    The opposite case comes up just as often. A model can be trained on a mix of feature store features and extra columns that were never in a feature table. In the documentation's example, browser is kept in the training DataFrame by not listing it in exclude_columns. The model's feature metadata covers only the feature-table features, so score_batch has nowhere to look up browser. The caller must supply it.

    How fe.score_batch treats each kind of column in the input DataFrame
    Column in batch_dfExampleWhat score_batch does
    Lookup key matching the feature table's primary keycustomer_idUses it to look up the packaged features from the feature table
    Column named the same as a packaged featureaccount_creation_dateUses the supplied values and doesn't look up that feature
    Training column that came from outside the feature storebrowserMust be present at inference; it is not looked up
    Added to the outputpredictionAdded next to the input and the looked-up feature columns

    Checkpoint 4 of 6· Check yourself

    A model was trained with df columns ['customer_id', 'browser', 'rating'], a FeatureLookup for total_purchases_30d, and exclude_columns=['customer_id']. Which batch_df is correct for fe.score_batch?

    Checkpoint 5 of 6· Exam question

    A team trained a scikit-learn model using a DataFrame produced by `TrainingSet.load_df()`, but at the end of the notebook they logged it using `mlflow.sklearn.log_model(model, "model")` instead of the Feature Engineering client's logging method. When they later call `fe.score_batch()` against a batch DataFrame containing only lookup keys, the call fails. What is the most likely cause?

    Sources2

    4.When a lookup key does not exist

    Batch data often includes entities the feature table has never seen, such as a customer who signed up after the last feature refresh. score_batch doesn't raise an error for these rows. The lookup returns a missing value, and the documentation states which missing value you get. For offline scoring with fe.score_batch, a missing feature comes back as NaN. With Model Serving it depends on the request: if none of the lookup keys exist, the value is None; if only some are missing, it is NaN. Because the same model may run in both places, the documentation tells you to write model code that handles both None and NaN.

    Checkpoint 6 of 6· Check yourself

    A nightly fe.score_batch job includes a customer_id that has no row in the feature table. What does the model receive for that customer's looked-up features?

    Sources2

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.The DataFrame passed to score_batch must already contain every feature column the model was trained on.Why is that wrong?

      Only the lookup keys, plus any columns that came from outside the feature store, are required. The packaged features are looked up automatically. If you include a column named like a feature, your values replace the lookup.

      Covered in Overriding features and supplying non-feature-store columns

    2. 2.Every column the model was trained on is looked up automatically at scoring time, including extra columns such as browser.Why is that wrong?

      Only features from a FeatureLookup are packaged as metadata. Columns that came straight from the training DataFrame must be supplied in the scoring DataFrame.

      Covered in Overriding features and supplying non-feature-store columns

    3. 3.If a lookup key isn't in the feature table, fe.score_batch fails.Why is that wrong?

      Scoring continues, and the missing feature values come back as NaN for offline score_batch. Model implementations are expected to handle missing values.

      Covered in When a lookup key does not exist

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “To package a model with feature metadata, use FeatureEngineeringClient.log_model (for Feature Engineering in Unity Catalog)”
      ↩︎ Why scoring needs only primary keys
      “In batch inference, feature values are retrieved from the offline store and joined with new data prior to scoring.”
      ↩︎ Why scoring needs only primary keys
      “For batch use cases, the model automatically retrieves the features it needs from Feature Store.”
      ↩︎ Why scoring needs only primary keys
      “The caller only needs to provide the primary key of the features used in the model”
      ↩︎ Key concept
    2. 2.
      “The DataFrame returned by score_batch() augments batch_df with”
      ↩︎ Calling fe.score_batch and reading its output
      “columns containing the feature values and a column containing model predictions.”
      ↩︎ Calling fe.score_batch and reading its output
      “To use custom feature values for scoring, include them in the DataFrame passed to FeatureEngineeringClient.score_batch”
      ↩︎ Overriding features and supplying non-feature-store columns
      “At inference, the DataFrame used in FeatureStoreClient.score_batch must include the browser column.”
      ↩︎ Overriding features and supplying non-feature-store columns
      “If none of the provided lookup keys exist, the value is None.”
      ↩︎ When a lookup key does not exist
      “Your model implementation should be able to handle both values.”
      ↩︎ When a lookup key does not exist
      “By default, a model packaged with feature metadata looks up features from feature tables at inference.”
      ↩︎ Exam trap 1
      “At inference, the DataFrame used in FeatureStoreClient.score_batch must include the browser column.”
      ↩︎ Exam trap 2
      “For offline applications using fe.score_batch, the returned value for a missing feature is NaN.”
      ↩︎ Exam trap 3
      “In this case the API looks up only the num_lifetime_purchases feature from Feature Store”
      ↩︎ Prediction
      “For offline applications using fe.score_batch, the returned value for a missing feature is NaN.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Scoring Feature Store Models: Time Series Tables, MLflow pyfunc and Model Serving

    Spotted a mistake, or was something unclear? Tell us.