What you will be able to do
- Pick the right client and table name for writing features in a Unity Catalog workspace
- Populate a new feature table with create_table and write_table, or straight from create_table
- Predict what mode='merge' does to rows that are in the DataFrame and rows that are not
- Keep a feature table fresh with a scheduled job or a streaming DataFrame
- Write correctly to a time series feature table
Key concept
Feature table in Unity Catalog — A feature table is an ordinary Delta table in Unity Catalog that has a primary key constraint. Writing features means putting rows into that table, keyed by the primary key, and the key is also what decides how new rows combine with the ones already there.
1.What you are writing to, and with which client
Databricks Feature Store gives you two ways to author features. With Feature Views, Databricks manages the pipelines for you. With feature tables, you write the values yourself. This lesson is about the second way, and that means you own the code that fills the table and keeps it current. In a workspace enabled for Unity Catalog, a feature table is not a special storage format. It is a Delta table with a primary key, named with the usual three levels: <catalog-name>.<schema-name>.<table-name>. Every example below writes to ml.recommender_system.customer_features.
Before you write anything, check which client you have. The current client is FeatureEngineeringClient, from the databricks-feature-engineering package. It comes pre-installed in Databricks Runtime 13.3 LTS ML and above. On a non-ML runtime you install it yourself with %pip install databricks-feature-engineering. The older FeatureStoreClient is for the legacy Workspace Feature Store only, and the databricks-feature-store package was deprecated as of version 0.17.0.
| Feature tables in | Python client | Package | Example table name |
|---|---|---|---|
| Unity Catalog | FeatureEngineeringClient | databricks-feature-engineering | ml.recommender_system.customer_features |
| Workspace Feature Store (legacy) | FeatureStoreClient | databricks-feature-engineering (0.2.0+) or deprecated databricks-feature-store | recommender_system.customer_features |
One more limit: the client runs only on Databricks. You can install it locally to mock its calls in unit tests, but it cannot call Feature Engineering APIs from your laptop.
Checkpoint 1 of 8· Check yourself
A data scientist in a Unity Catalog-enabled workspace needs to write features to ml.recommender_system.customer_features from a notebook. Which client should they instantiate?
Feature tables in Unity Catalog use FeatureEngineeringClient. FeatureStoreClient is for the legacy Workspace Feature Store.
“To work with feature tables in Unity Catalog, use FeatureEngineeringClient. To use Workspace Feature Store, you must use FeatureStoreClient.”Source: docs.databricks.com
2.The first write: populating a new feature table
The Python workflow has three steps. First, write functions that compute the features. Each one should return a Spark DataFrame with a unique primary key, which can be one column or several. Second, create the table with create_table. Third, populate it with write_table. The DataFrame comes first because create_table can take its schema straight from it.
from databricks.feature_engineering import FeatureEngineeringClient
fe = FeatureEngineeringClient()
# Prepare feature DataFrame
def compute_customer_features(data):
''' Feature computation code returns a DataFrame with 'customer_id' as primary key'''
pass
customer_features_df = compute_customer_features(df)
# Create feature table with `customer_id` as the primary key.
# Take schema from DataFrame output by compute_customer_features
customer_feature_table = fe.create_table(
name='ml.recommender_system.customer_features',
primary_keys='customer_id',
schema=customer_features_df.schema,
description='Customer features'
)That call creates an empty table. You can skip the separate write by passing the DataFrame to create_table as df=customer_features_df. The docs say that this "automatically saves the features to the underlying Delta table." For a composite key, pass a list such as primary_keys=['customer_id', 'date'].
You don't have to use Python. If you create the table in Databricks SQL with a PRIMARY KEY constraint, you can then write to it like any other Delta table. A materialized view or streaming table published by Lakeflow pipelines with a primary key is also a feature table, and you write to it like any other pipeline dataset. An existing Delta table becomes writable as a feature table once it has a primary key. Only the table owner can add that constraint, using ALTER TABLE ... ADD CONSTRAINT ... PRIMARY KEY.
Checkpoint 2 of 8· Put it in order
Put the steps for creating and populating a feature table with the Python client in order.
- 1.Instantiate FeatureEngineeringClient and call create_table
- 2.Populate the feature table using write_table
- 3.Write functions that compute the features and return a Spark DataFrame with a unique primary key
create_table takes its schema from the computed DataFrame, so the features have to be computed first. write_table then fills the table.
“Populate the feature table using write_table.”Source: docs.databricks.com
No. A SQL-created table with a primary key can be written like any other Delta table and still works as a feature table. write_table is the Python client's route, not the only one.
Sources2
3.Updating with mode='merge'
After the first load, most writes are updates. The Unity Catalog docs describe two kinds: adding new features, and changing specific rows based on the primary key. Both go through write_table with mode='merge'.
Merge works key by key. Rows whose keys are in the DataFrame get updated. Keys the table doesn't have yet come in as new rows. Rows whose keys aren't in the DataFrame are left alone. That's why a merge write is safe to run on a schedule with only the latest batch of entities.
from databricks.feature_engineering import FeatureEngineeringClient
fe = FeatureEngineeringClient()
customer_features_df = compute_customer_features(data)
fe.write_table(
df=customer_features_df,
name='ml.recommender_system.customer_features',
mode='merge'
)Merge is also how you add new feature columns. One way is to extend the existing computation function and write its output. That updates the table schema and merges the new values by primary key. The other way is a separate function that returns only the new features. Its DataFrame must still contain the table's primary key, plus the partition keys if any are defined, so each value lands on the right row. Some metadata should not change through writes: the primary key, the partition key, and the name or data type of an existing feature. Changing any of these breaks the downstream training and serving pipelines that read the table.
| mode | Where the docs show it | Effect on existing rows |
|---|---|---|
| 'merge' | FeatureEngineeringClient.write_table (Unity Catalog) and FeatureStoreClient.write_table | Rows with matching keys are updated. Rows whose keys are not in the DataFrame remain unchanged. |
| 'overwrite' | FeatureStoreClient.write_table (legacy Workspace Feature Store example) | A full refresh of the feature table |
Checkpoint 3 of 8· Fill the gap
This unit test checks that a feature update function updates rows by key and leaves other rows alone. Which value completes the asserted call?
mock_write_table.assert_called_once_with(
name='ml.recommender_system.customer_features',
df=customer_features_df,
mode=' ? '
)The documented update call for feature tables uses mode='merge', which updates rows by primary key and leaves other rows unchanged.
Source: docs.databricks.comCheckpoint 4 of 8· Exam question
A feature engineering team maintains a Unity Catalog feature table `customer_features` with primary key `customer_id`. Each night, a batch job computes updated values for only the customers who transacted that day, and the team wants those rows inserted or updated while every other customer's existing feature values stay exactly as they are. Which write call achieves this without touching unrelated rows?
Correct answer: B — Call `fe.write_table(name="customer_features", df=nightly_df, mode="merge")`, since merge mode inserts new primary keys and updates matching ones while leaving unmatched rows unchanged.
- A. This description of overwrite mode is incorrect: overwrite mode replaces the entire table's schema and data with the incoming DataFrame rather than selectively updating only the rows it contains. Using it here would delete every customer not present in the nightly batch.
- B. Merge mode is correct here because it performs an upsert keyed on the primary key column: rows whose key matches an existing row get updated, rows with a new key get inserted, and rows whose key is absent from the incoming DataFrame are left untouched.
- C. `create_table` is meant to define a brand-new feature table's schema and primary keys, not to update an existing one; calling it again against a table that already exists conflicts with its purpose rather than performing a targeted upsert.
- D. Writing through the raw Spark DataFrame API bypasses the feature engineering client entirely, so primary key enforcement and feature table metadata are not applied, and a plain append can create duplicate rows for the same primary key instead of upserting them.
4.Keeping features fresh: scheduled jobs and streaming writes
A feature table is only useful while its values are current. Databricks recommends a job that runs your update notebook on a schedule, for example daily. If you already have an unscheduled job, you can convert it to a scheduled one. The notebook just runs the merge write from the previous section.
Batch isn't the only option, because write_table also accepts a streaming DataFrame. In that case it starts a streaming feature pipeline and returns a StreamingQuery object instead of finishing as a one-off write. The call looks the same and still uses merge mode.
customer_transactions = spark.readStream.table("prod.events.customer_transactions")
stream_df = compute_additional_customer_features(customer_transactions)
fe.write_table(
df=stream_df,
name='ml.recommender_system.customer_features',
mode='merge'
)Each write is also recorded. Feature table metadata tracks the data sources a table was built from, and the notebooks and jobs that created it or wrote to it. That record shows where a feature value came from.
Checkpoint 5 of 8· Check yourself
What does write_table return when you pass it a streaming DataFrame?
Passing a streaming DataFrame to write_table creates a streaming feature computation pipeline and returns a StreamingQuery.
“pass a streaming DataFrame as an argument to write_table. This method returns a StreamingQuery object.”Source: docs.databricks.com
Checkpoint 6 of 8· Exam question
A data scientist runs `fe.write_table(name="loan_features", df=updates_df, mode="merge")` against a feature table whose primary key is the composite `(loan_id, as_of_date)`, and the call fails during validation. `updates_df` contains `loan_id`, `credit_score`, and `utilization_ratio`, but no `as_of_date` column. What is the most likely cause of the failure?
Correct answer: D — The DataFrame is missing one of the feature table's primary key columns, and `write_table` requires every primary key column to be present so it can determine which rows to update or insert.
- A. This claim about merge mode is invented: merge mode works the same way regardless of whether the primary key is a single column or a composite key, so restricting composite-key tables to overwrite mode is not an actual constraint of the API.
- B. Column order in the incoming DataFrame does not matter to `write_table`; it matches columns by name against the feature table's schema, so this is not the reason the call fails.
- C. Publishing to an online store does not place a lock on offline writes to the feature table, so a prior online sync would not block this `write_table` call.
- D. Because the feature table's primary key is `(loan_id, as_of_date)`, every write must supply both columns so the client can identify which existing rows to update and which new rows to insert; omitting `as_of_date` leaves the operation unable to resolve the primary key, which is why validation fails.
5.Writing to time series feature tables
A time series feature table includes a timestamp column in its primary key and marks it with TIMESERIES in SQL, or with timeseries_columns in create_table. You still write to it with write_table and merge mode. There is one rule that ordinary feature tables don't have.
For time series tables, every write must include values for all of the table's features. A DataFrame with just one or two feature columns won't do. The docs say this rule keeps feature values from becoming sparse across timestamps. Streaming writes to time series tables are supported.
fe = FeatureEngineeringClient()
# daily_users_batch_df DataFrame contains the following columns:
# - user_id
# - ts
# - purchases_30d
# - is_free_trial_active
fe.write_table(
"ml.ads_team.user_features",
daily_users_batch_df,
mode="merge"
)Checkpoint 7 of 8· Check yourself
ml.ads_team.user_features is a time series feature table with features purchases_30d and is_free_trial_active. A job computes only purchases_30d and merges a DataFrame of (user_id, ts, purchases_30d). How does that compare with the documented requirement?
Unlike ordinary feature tables, time series feature tables need every write to supply all features. Streaming writes to them are supported, so the last option is wrong.
“your DataFrame must supply values for all features of the feature table, unlike regular feature tables.”Source: docs.databricks.com
Checkpoint 8 of 8· Exam question
A team rebuilds a `product_features` feature table from scratch every Sunday by recomputing every feature for every product from raw source tables, and the resulting DataFrame should completely replace the previous week's contents, including dropping any columns that no longer exist. Which write approach is correct for this weekly job?
Correct answer: A — Call `fe.write_table(name="product_features", df=full_df, mode="overwrite")`, since overwrite mode replaces both the schema and the data of the table with the incoming DataFrame.
- A. Overwrite mode is correct for a full weekly rebuild because it replaces the table's schema as well as its data with whatever the incoming DataFrame contains, including dropping columns that are no longer present, which matches exactly what this job needs.
- B. Merge mode does not drop columns; it only upserts rows by primary key while leaving the existing schema and any untouched rows in place, so it would not remove obsolete columns as this weekly rebuild requires.
- C. `create_table` is meant to define a new feature table once, and calling it repeatedly against a table that already exists conflicts with an existing table rather than performing a clean weekly refresh of its contents.
- D. Dropping the Delta table with raw SQL destroys the feature table's Unity Catalog registration and metadata outside of the feature engineering client, and merge mode afterward cannot restore metadata it has no record of, so this approach breaks the table rather than rebuilding it correctly.
Sources6
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.write_table with mode='merge' makes the table match the DataFrame, so keys missing from the DataFrame get deleted.Why is that wrong?
Merge updates and inserts rows by primary key. Rows whose keys aren't in the DataFrame are left alone.
Covered in Updating with mode='merge'
2.FeatureStoreClient.write_table is the way to write to feature tables in Unity Catalog.Why is that wrong?
Unity Catalog feature tables use FeatureEngineeringClient. FeatureStoreClient is for the legacy Workspace Feature Store.
3.Like an ordinary feature table, a time series feature table accepts a merge with only a subset of its feature columns.Why is that wrong?
Writes to time series feature tables must supply values for every feature, which keeps values from becoming sparse across timestamps.
Covered in Writing to time series feature tables
4.Since write_table can change the table schema, it's fine to rename a feature or change its data type in a later write.Why is that wrong?
Feature names, data types, the primary key and the partition key should not change. Changing them breaks downstream training and serving pipelines.
Covered in Updating with mode='merge'
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Feature tables, which you populate yourself by writing feature values to a Delta table.”
↩︎ What you are writing to, and with which client - 2.
“Feature tables, like other data assets in Unity Catalog, are accessed using a three-level namespace”
↩︎ What you are writing to, and with which client“is pre-installed in Databricks Runtime 13.3 LTS ML and above”
↩︎ What you are writing to, and with which client“The output of each function should be an Apache Spark DataFrame with a unique primary key.”
↩︎ The first write: populating a new feature table“This code automatically saves the features to the underlying Delta table.”
↩︎ The first write: populating a new feature table“After the table is created, you can write data to it like other Delta tables, and it can be used as a feature table.”
↩︎ The first write: populating a new feature table“Only the table owner can declare primary key constraints.”
↩︎ The first write: populating a new feature table“Rows whose primary key does not exist in the DataFrame sent in the write_table call remain unchanged.”
↩︎ Updating with mode='merge'“This updates the feature table schema and merges new feature values based on the primary key.”
↩︎ Updating with mode='merge'“The DataFrame returned by this new computation function must contain the feature tables' primary and partition keys (if defined).”
↩︎ Updating with mode='merge'“create a job that runs a notebook to update your feature table on a regular basis, such as every day”
↩︎ Keeping features fresh: scheduled jobs and streaming writes“pass a streaming DataFrame as an argument to write_table. This method returns a StreamingQuery object.”
↩︎ Keeping features fresh: scheduled jobs and streaming writes“In Unity Catalog, any Delta table with a primary key constraint can serve as a feature table.”
↩︎ Key concept“Rows whose primary key does not exist in the DataFrame sent in the write_table call remain unchanged.”
↩︎ Exam trap 1“Altering them will cause downstream pipelines that use features for training and serving models to break.”
↩︎ Exam trap 4“Populate the feature table using write_table.”
↩︎ Checkpoint - 3.
“As of version 0.17.0, databricks-feature-store has been deprecated.”
↩︎ What you are writing to, and with which client“It does not support calling Feature Engineering in Unity Catalog or Feature Store APIs from a local environment”
↩︎ What you are writing to, and with which client“To work with feature tables in Unity Catalog, use FeatureEngineeringClient. To use Workspace Feature Store, you must use FeatureStoreClient.”
↩︎ Exam trap 2“To work with feature tables in Unity Catalog, use FeatureEngineeringClient. To use Workspace Feature Store, you must use FeatureStoreClient.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/machine-learning/feature-store/workspace-feature-store/feature-tablesOfficial docs
“Overwrite mode does a full refresh of the feature table”
↩︎ Updating with mode='merge' - 5.
“tracks the data sources from which a table was generated and the notebooks and jobs that created or wrote to the table”
↩︎ Keeping features fresh: scheduled jobs and streaming writes - 6.
“your DataFrame must supply values for all features of the feature table, unlike regular feature tables.”
↩︎ Writing to time series feature tables“Streaming writes to time series feature tables is supported.”
↩︎ Writing to time series feature tables“your DataFrame must supply values for all features of the feature table, unlike regular feature tables.”
↩︎ Exam trap 3