What you will be able to do
- Prepare a scoring DataFrame for a model trained on time series feature tables
- Choose between fe.score_batch and mlflow.pyfunc for batch inference, and know when pyfunc isn't supported
- Explain how Model Serving looks up features from an online store and how to override them in a request
- Pick FeatureEngineeringClient over the deprecated FeatureStoreClient for Unity Catalog feature tables
1.Scoring with time series feature tables
When feature values change over time, a model is often trained on time series feature tables. These tables have a timestamp key, and each training row gets the latest feature values known as of that row's timestamp. Scoring has to follow the same rule, otherwise predictions would use values the model never saw paired with a timestamp during training. score_batch does this for you. It uses the metadata stored with the model to run point-in-time lookups at inference. To do that, it needs a timestamp for each row of your DataFrame.
# 5. Batch scoring with automatic feature lookup
# inference_df must contain the same entity and timeseries columns
# used during training. Features are automatically computed.
predictions = fe.score_batch(
model_uri="models:/main.ecommerce.fraud_model/1",
df=inference_df,
)
predictions.display()The timestamp column must match exactly. It needs the same name and the same DataType as the timestamp_lookup_key used in the training FeatureLookup. A lookback window set during training also applies to batch inference. For performance, when Photon is enabled you can pass use_spark_native_join=True to FeatureEngineeringClient.score_batch (this requires databricks-feature-engineering 0.6.0 or above). The example also shows a Unity Catalog registry URI with a three-level name, models:/main.ecommerce.fraud_model/1.
Checkpoint 1 of 5· Check yourself
A model was trained with a FeatureLookup whose timestamp_lookup_key is ts (TimestampType). Your scoring DataFrame has the primary key and a column event_time of the same type. What must you change before calling fe.score_batch?
The scoring DataFrame needs a timestamp column with the same name and DataType as the training timestamp_lookup_key. use_spark_native_join only speeds up the lookup when Photon is enabled.
“must contain a timestamp column with the same name and DataType as the timestamp_lookup_key of the FeatureLookup”Source: docs.databricks.com
2.Batch inference through mlflow.pyfunc
score_batch isn't the only batch scoring path. Models logged with the client's log_model also work with the MLflow pyfunc interface. You load the model with mlflow.pyfunc.load_model and call predict. The model then retrieves feature values from Feature Store and joins in any values you supplied, just as score_batch does. You still provide the primary keys. There are two limits. This path requires MLflow 2.11 or above, and it doesn't support models that use time series feature tables.
# batch_df has columns 'customer_id' and 'product_id'
model = mlflow.pyfunc.load_model(model_version_uri)
# If result_type parameter is provided in log_model
predictions = model.predict(df, {"result_type":"double"})
# If result_type parameter is NOT provided in log_model
model._model_impl.set_result_type("double")
predictions = model.predict(df)You can pass the result_type parameter to predict only if a default for it was set in log_model. Databricks Runtime for ML 15.0 and above sets that default automatically. For models logged on earlier runtimes, call set_result_type before predict, as the second branch of the sample shows.
Checkpoint 2 of 5· Check yourself
Which statement about batch inference with mlflow.pyfunc on a feature store model is correct?
pyfunc batch inference requires MLflow 2.11 or above, looks up features using the primary keys you supply, and doesn't support time series feature tables.
“Batch inference with MLflow requires MLflow version 2.11 and above.”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
A fraud detection model was trained using features from a Unity Catalog feature table, including a `transaction_velocity` feature. Before calling `fe.score_batch()`, an engineer added their own precomputed `transaction_velocity` column to the batch DataFrame using a faster streaming aggregation rather than relying on the feature table. What happens when `score_batch()` runs?
Correct answer: A — `score_batch()` uses the precomputed `transaction_velocity` values already present in the batch DataFrame for those rows and only retrieves the remaining features that are missing from the feature table.
- A. When a feature column is already present in the DataFrame passed to `score_batch()`, the API treats that as an override and uses the supplied value instead of looking it up, while still fetching any other declared features that are absent.
- B. Supplying a feature column that matches a declared feature name does not trigger a validation error; it is the documented way to override a specific feature value rather than a schema conflict.
- C. The API does not discard user-supplied feature values in favor of the feature table; doing so would make it impossible to override a feature for testing or correction, which the scoring API explicitly supports.
- D. There is no struct-merging behavior for overlapping feature columns; `score_batch()` uses the existing column values directly rather than combining them with looked-up values into a nested type.
Sources3
3.Real-time scoring with Model Serving
The same packaged model can also be scored in real time. Batch scoring reads the offline Delta tables. For real-time scoring you publish the feature tables to an online store, such as a Databricks Online Feature Store or Amazon DynamoDB, and serve the model with Model Serving. Every feature the model was logged with is then looked up from the online store when a request arrives. You can publish the tables at any time before deployment, even after training. Overriding a feature works much like it does in batch: include the feature values in the API payload, using the data type the model expects. Time series tables behave differently online. The lookback window applies only to training and batch inference, and online inference always uses the latest feature value.
| Scoring path | Where features come from | Constraint from the docs |
|---|---|---|
| fe.score_batch | Feature tables in the offline store, joined to the input DataFrame | Time series models need a matching timestamp column |
| mlflow.pyfunc.load_model + predict | Feature Store, joined with values supplied at inference | MLflow 2.11 or above; time series feature tables not supported |
| Model Serving | Features published to online stores | Override features by including them in the request payload |
Checkpoint 4 of 5· Match them up
Match each scoring situation to how features are handled
Tap a term, then the definition that fits it.
Model Serving uses the online store and accepts overrides in the payload. Batch scoring joins from offline tables. Online lookups on time series tables ignore the lookback window.
“are automatically looked up from online stores for model scoring.”Source: docs.databricks.com
4.Which client to call
Older examples call FeatureStoreClient.score_batch. That client belongs to the legacy Workspace Feature Store, which is deprecated and available only in workspaces created before August 19, 2024. For feature tables in Unity Catalog, use FeatureEngineeringClient from the databricks-feature-engineering package, on any Databricks Runtime ML version. The databricks-feature-store package was deprecated as of version 0.17.0, and its modules now ship in databricks-feature-engineering. The scoring call has the same shape in both clients, so the main differences are the package you import and how the tables are named: Unity Catalog tables use three-level names such as ml.recommender_system.customer_features.
Checkpoint 5 of 5· Check yourself
You are scoring a model trained on Unity Catalog feature tables. Which statement is correct?
The legacy Workspace Feature Store is deprecated, and Databricks recommends Feature Engineering in Unity Catalog, which uses FeatureEngineeringClient.
“Workspace Feature Store is deprecated. Databricks recommends using Feature Engineering in Unity Catalog.”Source: docs.databricks.com
FeatureEngineeringClient from databricks-feature-engineering, calling fe.score_batch(model_uri=..., df=...). FeatureStoreClient is for Workspace feature tables.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.mlflow.pyfunc.predict can batch-score any model logged with the Feature Engineering client, so score_batch is never needed.Why is that wrong?
pyfunc batch inference doesn't support models that use time series feature tables. Those must be scored with score_batch.
Covered in Batch inference through mlflow.pyfunc
2.FeatureStoreClient.score_batch is the current way to score models that use Unity Catalog feature tables.Why is that wrong?
Workspace Feature Store and its client are deprecated. Feature tables in Unity Catalog use FeatureEngineeringClient.
Covered in Which client to call
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Databricks Feature Store retrieves the appropriate features using point-in-time lookups with metadata packaged with the model during training.”
↩︎ Scoring with time series feature tables“For faster lookup performance when Photon is enabled, pass use_spark_native_join=True to FeatureEngineeringClient.score_batch.”
↩︎ Scoring with time series feature tables“The lookback window is applied during training and batch inference.”
↩︎ Scoring with time series feature tables“During online inference, the latest feature value is always used, regardless of the lookback window.”
↩︎ Real-time scoring with Model Serving“must contain a timestamp column with the same name and DataType as the timestamp_lookup_key of the FeatureLookup”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/machine-learning/feature-store/train-with-declarative-featuresOfficial docs
“score_batch uses the feature metadata stored with the model to automatically compute point-in-time correct features for inference, ensuring consistency with training.”
↩︎ Scoring with time series feature tables - 3.https://docs.databricks.com/aws/en/machine-learning/feature-store/train-models-with-feature-storeOfficial docs
“mlflow.pyfunc.predict retrieves feature values from Feature Store and also joins any values provided at inference time.”
↩︎ Batch inference through mlflow.pyfunc“You must provide the primary key(s) of the features used in the model.”
↩︎ Batch inference through mlflow.pyfunc“the result_type parameter can only be used if a default value”
↩︎ Batch inference through mlflow.pyfunc“Models that use time series feature tables are not supported.”
↩︎ Exam trap 1“To do batch inference with time series feature tables, use score_batch.”
↩︎ Prediction“Batch inference with MLflow requires MLflow version 2.11 and above.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/machine-learning/feature-store/automatic-feature-lookupOfficial docs
“include the feature values as a part of the API payload.”
↩︎ Real-time scoring with Model Serving“You can publish the feature table at any time prior to model deployment, including after model training.”
↩︎ Real-time scoring with Model Serving“are automatically looked up from online stores for model scoring.”
↩︎ Checkpoint - 5.
“As of version 0.17.0, databricks-feature-store has been deprecated.”
↩︎ Which client to call - 6.https://docs.databricks.com/aws/en/machine-learning/feature-store/workspace-feature-storeOfficial docs
“Workspace Feature Store is deprecated. Databricks recommends using Feature Engineering in Unity Catalog.”
↩︎ Which client to call“Workspace Feature Store is deprecated. Databricks recommends using Feature Engineering in Unity Catalog.”
↩︎ Exam trap 2