What you will be able to do
- Create a time series feature table with the TIMESERIES keyword or timeseries_columns
- Explain what goes wrong when a timestamp is a primary key but not declared as a time series column
- Turn an existing Delta table into a feature table with ALTER TABLE
- Declare a primary key on a Lakeflow pipelines materialized view and confirm the table in the Features UI
1.Time series feature tables
When feature values change over time, the feature table needs a time column in its primary key. That column must also be marked as the time series key, which lets training data be joined as of each label's timestamp rather than only on an exact match. Unity Catalog treats any table with a TIMESERIES primary key as a time series feature table.
In SQL, add the time column to the primary key and put the TIMESERIES keyword after it. This keyword needs Databricks Runtime 13.3 LTS or above.
CREATE TABLE ml.recommender_system.customer_features (
customer_id int NOT NULL,
ts timestamp NOT NULL,
feat1 long,
feat2 varchar(100),
CONSTRAINT customer_features_pk PRIMARY KEY (customer_id, ts TIMESERIES)
);In Python, list the time column in primary_keys and also name it in timeseries_columns:
Checkpoint 1 of 6· Fill the gap
Which argument turns on point-in-time logic for event_timestamp?
fe.create_table(
name="catalog.schema.user_behavior_features",
primary_keys=["user_id", "event_timestamp"],
? ="event_timestamp", # Enables point-in-time logic
df=features_df # DataFrame must contain primary keys and time series columns
)Feature Engineering in Unity Catalog uses timeseries_columns. timestamp_keys is the argument for the legacy Workspace Feature Store.
Source: docs.databricks.comPutting a date in the key without declaring it causes a second problem as well. If a DATE or TIMESTAMP primary key column is not declared as a time series column, you can't use the table with create_feature_spec(), create_training_set() or publish_table(). If what you really want is exact-match lookups on a date value, the documentation says to change the column type to STRING.
Time series tables also have their own rules:
- A time series table must have exactly one timestamp key and cannot have partition columns. - The timestamp key must be of TimestampType or DateType. - Databricks recommends no more than two primary key columns, for performant writes and lookups. - When you write to the table, your DataFrame must supply values for every feature. Regular feature tables don't require this. - Streaming writes to time series feature tables are supported.
Checkpoint 2 of 6· Check yourself
Which design is valid for a time series feature table in Unity Catalog?
A time series table needs exactly one timestamp key of TimestampType or DateType and no partition columns.
“A time series feature table must have one timestamp key and cannot have any partition columns.”Source: docs.databricks.com
Checkpoint 3 of 6· Exam question
An ML engineer is writing the Python call to create a new feature table for a churn model and needs to specify where the table will live. Which value should be passed as the `name` argument to `FeatureEngineeringClient.create_table` for it to be stored correctly in Unity Catalog?
Correct answer: A — A three-level identifier in `catalog.schema.table` form, like `ml_prod.churn_features.customer_churn`, matching Unity Catalog object naming.
- A. Correct. Unity Catalog feature tables are addressed with the same three-level namespace as every other Unity Catalog table, so `create_table` expects `catalog.schema.table`.
- B. Incorrect. Unity Catalog does not silently substitute a default catalog for a missing qualifier; the catalog must be stated explicitly as part of the three-level name.
- C. Incorrect. There is no dedicated system catalog that feature tables fall back to; each feature table is created in whichever catalog and schema the caller specifies.
- D. Incorrect. Feature tables are governed Delta tables registered in Unity Catalog, not raw files addressed by a DBFS mount path.
2.Turning an existing Delta table into a feature table
Often the features are already in a Delta table in Unity Catalog. You don't need to copy them. If the table already has a primary key constraint, it's already a feature table. If it doesn't, you add the constraint with ALTER TABLE DDL. Before you start, check that its column types are ones Feature Engineering supports (see the last section).
Only the table owner can declare primary key constraints. The owner's name is shown on the table detail page in Catalog Explorer, so check it before you try the DDL. There are two steps. First, set every primary key column to NOT NULL:
ALTER TABLE <full_table_name> ALTER COLUMN <pk_col_name> SET NOT NULLSecond, add the named constraint. To make the table a time series feature table at the same time, add TIMESERIES to one of the key columns, as in PRIMARY KEY(pk_col1 TIMESERIES, pk_col2, ...).
ALTER TABLE <full_table_name> ADD CONSTRAINT <pk_name> PRIMARY KEY(pk_col1, pk_col2, ...)Checkpoint 4 of 6· Put it in order
Put the steps for turning an existing Unity Catalog Delta table into a feature table in order
- 1.Find the table in the Features UI and use it as a feature table
- 2.Run ALTER COLUMN ... SET NOT NULL on each primary key column
- 3.Run ADD CONSTRAINT ... PRIMARY KEY on the table
- 4.Confirm you are the table owner, since only the owner can declare primary key constraints
Primary key columns must be NOT NULL before the constraint can be added, and the table shows up in the Features UI only once the constraint exists.
“the table appears in the Features UI and you can use it as a feature table.”Source: docs.databricks.com
Sources2
3.Feature tables from Lakeflow pipelines
The third creation path is Lakeflow pipelines. Any table a pipeline publishes with a primary key constraint can be used as a feature table. You declare the constraint in the dataset's own schema, either in Databricks SQL or in the Python @dp.materialized_view(schema=...) decorator. Support for table constraints in Lakeflow pipelines is in Public Preview, and the examples must run on the Lakeflow pipelines preview channel.
You can't bolt a constraint onto a pipeline-published streaming table or materialized view with ALTER TABLE the way you can with an ordinary Delta table. You change the schema in the dataset definition to include the primary key, then refresh the streaming table or materialized view.
CREATE MATERIALIZED VIEW customer_features (
customer_id int NOT NULL,
feat1 long,
feat2 varchar(100),
CONSTRAINT customer_features_pk PRIMARY KEY (customer_id)
) AS SELECT * FROM ...;Checkpoint 5 of 6· Match them up
Match each starting point to how you give it the primary key it needs
Tap a term, then the definition that fits it.
Every path ends in the same place, a Unity Catalog table with a primary key. Only a pipeline-published dataset needs its definition changed and refreshed rather than an ALTER TABLE.
“Any table published from Lakeflow pipelines that includes a primary key constraint can be used as a feature table.”Source: docs.databricks.com
Sources2
4.Confirming the table in the Features UI and checking data types
Whichever path you used, you can confirm the result in the Features UI. Click Features in the sidebar and pick a catalog. Every feature table in that catalog is listed with its owner, the online stores it has been published to, the last time a notebook or job wrote to it, its key-value tags and its text comments. You can search by table name, feature, comment or tag. If a table you expected is missing, it has no primary key, so go back to the ALTER TABLE steps.
The table also gets the usual Unity Catalog benefits. Its source data is recorded as lineage, access is governed by Unity Catalog, and it is available in any workspace that has access to the catalog.
Check column types before you create or convert a table. The supported PySpark types are IntegerType, FloatType, BooleanType, StringType, DoubleType, LongType, TimestampType, DateType, ShortType, ArrayType, BinaryType, DecimalType, MapType and StructType. StructType needs Feature Engineering v0.6.0 or above. These types map onto common ML feature shapes:
| Feature shape | PySpark type |
|---|---|
| Dense vectors, tensors, embeddings | ArrayType |
| Sparse vectors, tensors, embeddings | MapType |
| Text | StringType |
Checkpoint 6 of 6· Check yourself
A Delta table in Unity Catalog doesn't appear in the Features UI, even though the user has access to its catalog. What is the most likely cause?
The Features UI lists every Unity Catalog table that has a primary key. A missing table points to a missing constraint.
“If you don't see a table on this page, see how to add a primary key constraint on the table.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Putting a timestamp column in the primary key is enough to get point-in-time joins.Why is that wrong?
Without TIMESERIES in SQL, or timeseries_columns in create_table, the column is matched exactly, and create_training_set rejects tables with undeclared DATE or TIMESTAMP key columns.
Covered in Time series feature tables
2.Anyone with write access to a Delta table can add the primary key constraint that makes it a feature table.Why is that wrong?
Only the table owner can declare primary key constraints. The owner's name is shown on the table detail page in Catalog Explorer.
Covered in Turning an existing Delta table into a feature table
Practise it for real
Create a Unity Catalog feature table end to end, then confirm that Databricks recognizes it
1.In a notebook on Databricks Runtime 13.3 LTS ML or above, run CREATE CATALOG IF NOT EXISTS <catalog-name> (or USE CATALOG for an existing one), then CREATE SCHEMA IF NOT EXISTS <schema-name>.
Why: Feature tables need a catalog and a schema, and creating each one takes its own privilege.
You should see: Both statements succeed. A privilege error means you are missing CREATE CATALOG, USE CATALOG or CREATE SCHEMA.
2.Build a Spark DataFrame of features with a unique key column, instantiate FeatureEngineeringClient, and call fe.create_table with a three-level name, primary_keys and df=.
Why: Passing df saves the features to the underlying Delta table as part of creating it.
You should see: A Delta table exists at catalog.schema.table containing your rows.
3.Compute a new batch of the same features and call fe.write_table(name, df, mode="merge").
Why: write_table is how a feature table is populated after it has been created.
You should see: The call succeeds and the table holds the new batch.
4.Click Features in the sidebar, select your catalog, and search for the table name.
Why: Any Unity Catalog table with a primary key is listed here automatically.
You should see: The table is listed with you as owner and a recent write time.
5.Create a second Delta table without a primary key, then run ALTER COLUMN ... SET NOT NULL and ADD CONSTRAINT ... PRIMARY KEY on it.
Why: This shows that an existing table becomes a feature table once it has the constraint.
You should see: The second table is missing from the Features UI before the ALTER and appears after it.
Stuck? Get a nudge
If the ALTER TABLE ADD CONSTRAINT step fails, check that every key column is NOT NULL and that you own the table.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“In Unity Catalog, any table with a TIMESERIES primary key is a time series feature table.”
↩︎ Time series feature tables“The timestamp key column must be of TimestampType or DateType.”
↩︎ Time series feature tables“your DataFrame must supply values for all features of the feature table, unlike regular feature tables.”
↩︎ Time series feature tables“change the column type to STRING instead.”
↩︎ Time series feature tables“These APIs require that all DATE and TIMESTAMP primary key columns are declared as timeseries columns.”
↩︎ Exam trap 1“it matches only rows with an exact time match instead of matching all rows prior to the timestamp.”
↩︎ Prediction“A time series feature table must have one timestamp key and cannot have any partition columns.”
↩︎ Checkpoint - 2.
“The TIMESERIES keyword requires Databricks Runtime 13.3 LTS or above.”
↩︎ Time series feature tables“If the table does not have a primary key defined, you must update the table using ALTER TABLE DDL statements to add the constraint.”
↩︎ Turning an existing Delta table into a feature table“Set primary key columns to NOT NULL.”
↩︎ Turning an existing Delta table into a feature table“Lakeflow pipelines support for table constraints is in Public Preview.”
↩︎ Feature tables from Lakeflow pipelines“Only the table owner can declare primary key constraints.”
↩︎ Exam trap 2“the table appears in the Features UI and you can use it as a feature table.”
↩︎ Checkpoint“Any table published from Lakeflow pipelines that includes a primary key constraint can be used as a feature table.”
↩︎ Checkpoint - 3.
“To access the Features UI, click Features in the sidebar.”
↩︎ Confirming the table in the Features UI and checking data types“Feature tables, functions, and models are automatically available in any workspace that has access to the catalog.”
↩︎ Confirming the table in the Features UI and checking data types“If you don't see a table on this page, see how to add a primary key constraint on the table.”
↩︎ Checkpoint - 4.
“You can store dense vectors, tensors, and embeddings as ArrayType.”
↩︎ Confirming the table in the Features UI and checking data types