What you will be able to do
- List the requirements an offline feature table must meet before it can be published online
- Explain how a time series designation changes what the online table contains
- Publish a feature table with publish_table and choose between TRIGGERED, CONTINUOUS and SNAPSHOT modes
- Query and correctly delete an online feature table
1.Preparing an offline table for publishing
An online feature table always comes from an offline one. Once a Databricks online store, created with create_online_store, reaches the AVAILABLE state, "The publish_table API synchronizes data from your offline feature table to online store created using create_online_store API." The offline Delta table has to meet three requirements first, whether or not it has a time series column:
- Primary key constraint: required for publishing to an online store. - Non-nullable primary keys: primary key columns cannot contain NULL values. - Change Data Feed enabled: required for the CONTINUOUS and TRIGGERED publish modes.
The last one matters because the sync is incremental. The publish pipeline reads changes from the offline table instead of copying the whole table every time. The docs give two DDL statements to get an existing table ready:
-- Enable CDF if not already enabled
ALTER TABLE catalog.schema.your_feature_table
SET TBLPROPERTIES ('delta.enableChangeDataFeed' = 'true');
-- Ensure primary key columns are not nullable
ALTER TABLE catalog.schema.your_feature_table
ALTER COLUMN user_id SET NOT NULL;Checkpoint 1 of 5· Fill the gap
Which table property must be enabled before a feature table can be published in TRIGGERED or CONTINUOUS mode?
-- Enable CDF if not already enabled
ALTER TABLE catalog.schema.your_feature_table
SET TBLPROPERTIES (' ? ' = 'true');TRIGGERED and CONTINUOUS publishing both read incremental changes from the offline table, which requires Delta Change Data Feed.
Source: docs.databricks.comSources1
2.Time series tables: the online copy keeps only the latest value
This is one of the clearest differences between the two kinds of table. Offline, a time series feature table supports point-in-time joins: create_training_set and score_batch do an as-of join so that each training row sees the feature values that were known at its timestamp. Online, a real-time endpoint usually wants the current value. So when the offline table has a time series designation, rows that share a primary key but have different time series values are deduplicated in the publish pipeline.
If an endpoint really needs historical values, create the offline table *without* a time series designation. Then "All rows from the source (offline) table are published without deduplication." Even then the online table does not do as-of queries: "This is an exact-match lookup on a string primary key, not a point-in-time (as-of) or temporal range query." To use a DATE or TIMESTAMP column as a plain lookup key, change the column type to STRING.
| Offline table created… | What is published online | Typical use |
|---|---|---|
| With time series designation | Latest feature values per entity ID (deduplicated in the publish pipeline) | Online model or feature serving endpoints |
| Without time series designation | All rows, without deduplication | Exact-match retrieval of a specific historical value by entity ID and date/timestamp, for example back-testing served predictions |
Checkpoint 2 of 5· Check yourself
A team publishes an offline table without a time series designation so that every historical row reaches the online store. Which kind of lookup can the online table serve?
Point-in-time (as-of) semantics exist offline. Online, even with every row published, the lookup is an exact match on the key.
“This is an exact-match lookup on a string primary key, not a point-in-time (as-of) or temporal range query.”Source: docs.databricks.com
3.publish_table and the three publish modes
Publishing is a single call on FeatureEngineeringClient. You pass the online store, the offline source_table_name and an online_table_name. The call creates the online table if it doesn't already exist, syncs the feature data from the offline table, and sets up the infrastructure that keeps the online store in sync with the offline table from then on.
from databricks.ml_features.entities.online_store import DatabricksOnlineStore
# Get the online store instance
# For Lakebase Autoscaling projects creating using the Lakebase API or UI,
# `name` is the last part of the resouce name: projects/{online_store_name}
online_store = fe.get_online_store(name="my-online-store")
# Publish the feature table to the online store
fe.publish_table(
online_store=online_store,
source_table_name="catalog_name.schema_name.feature_table_name",
# for online_table_name, the catalog name, schema name, and table name each are limited to a maximum of 63 bytes
online_table_name="catalog_name.schema_name.online_feature_table_name",
# `publish_mode` argument is optional and defaults to "TRIGGERED" mode if not specified
)How fresh the online copy is depends on publish_mode. This parameter replaced the older streaming parameter, and passing streaming=True is equivalent to publish_mode="CONTINUOUS".
| publish_mode | How the online table is updated | Needs Change Data Feed? |
|---|---|---|
| TRIGGERED (default) | Incremental updates from the offline table, run through the API or on a schedule (for example a scheduled Lakeflow Job that runs publish_table) | Yes |
| CONTINUOUS | A streaming pipeline updates the online store immediately as new data is written to the offline table | Yes |
| SNAPSHOT | One-time sync that copies all data from the source table; efficient when many existing rows change between syncs | Not listed as a requirement |
Checkpoint 3 of 5· Check yourself
You call publish_table without a publish_mode argument. How will the online table be kept up to date?
TRIGGERED is the default mode. It applies incremental changes from the offline table each time the sync runs, whether you start it through the API or on a schedule.
“Default. Incrementally updates the online table with changes from the offline table using the API or on a schedule.”Source: docs.databricks.com
Checkpoint 4 of 5· Exam question
An online store already serves a fraud-scoring endpoint, but the team notices the endpoint keeps scoring against feature values that are several hours stale relative to the offline feature table, which is updated continuously as transactions stream in. Which change to the publish configuration most directly reduces this staleness to a matter of seconds?
Correct answer: C — Switch the publish mode to CONTINUOUS so a streaming pipeline propagates each offline table change to the online store within seconds of it landing.
- A. SNAPSHOT performs a one-time full copy at the moment it runs and does not keep syncing afterward, so it does not reduce ongoing staleness between refreshes.
- B. TRIGGERED mode syncs incrementally on the cadence it is scheduled at; running it nightly reproduces the multi-hour staleness rather than reducing it to seconds.
- C. CONTINUOUS mode runs a streaming pipeline that propagates each change from the offline table to the online store almost immediately, which is what brings staleness down to seconds.
- D. Change Data Feed is what lets the publish job detect incremental changes at all; disabling it removes the mechanism TRIGGERED and CONTINUOUS modes depend on instead of speeding anything up.
Sources1
4.Querying and deleting online tables
Once the published table shows as AVAILABLE, you can work with it in two ways, neither of which looks like the offline Delta table. In the Unity Catalog UI you can open the online table to see sample data and check its schema. In the SQL editor you can run PostgreSQL queries against online feature tables, because the store is Lakebase. To serve the features to applications you create a Feature Serving endpoint. Models trained on Databricks features already carry lineage to those features, and when they're deployed they use Unity Catalog to find the matching online tables.
Deleting is where people most often treat an online table like an ordinary table. The only recommended way is the Databricks SDK, which removes the table from both Unity Catalog and the database. A SQL DROP TABLE, or the SDK command that deletes a synced table, leaves the data in the underlying database storage. Before you delete, make sure no model serving or feature serving endpoint still uses the online features, because removing a published table can break downstream dependencies.
Checkpoint 5 of 5· Fill the gap
Which method fully removes an online feature table from both Unity Catalog and the online database?
from databricks.sdk import WorkspaceClient
w = WorkspaceClient()
w.feature_store. ? (online_table_name="catalog_name.schema_name.online_feature_table_name")w.feature_store.delete_online_table is the only recommended method. It deletes the table from Unity Catalog and from the database together.
Source: docs.databricks.comSources1
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Publishing a time series feature table online keeps its full history, so real-time endpoints get the same point-in-time lookups as training.Why is that wrong?
With a time series designation, the publish pipeline deduplicates rows and keeps only the latest values per entity. Without one, every row is published, but lookups are exact matches and never as-of queries.
Covered in Time series tables: the online copy keeps only the latest value
2.Running DROP TABLE on an online feature table removes it completely.Why is that wrong?
DROP TABLE leaves the data in the underlying database storage. Only the SDK's delete_online_table removes the table from both Unity Catalog and the database.
Covered in Querying and deleting online tables
3.A Delta feature table with a primary key can be published in the default mode with no further changes.Why is that wrong?
The default TRIGGERED mode, like CONTINUOUS, requires Change Data Feed on the offline table, and primary key columns must also be non-nullable.
Covered in Preparing an offline table for publishing
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The publish_table API synchronizes data from your offline feature table to online store created using create_online_store API.”
↩︎ Preparing an offline table for publishing“Non-nullable primary keys: Primary key columns cannot contain NULL values”
↩︎ Preparing an offline table for publishing“All rows from the source (offline) table are published without deduplication.”
↩︎ Time series tables: the online copy keeps only the latest value“Set up the necessary infrastructure for keeping the online store in sync with the offline table.”
↩︎ publish_table and the three publish modes“set up with a streaming pipeline to immediately update the online store as new data written to the offline feature table”
↩︎ publish_table and the three publish modes“Performs a one-time sync that copies all data from the source table to the online store.”
↩︎ publish_table and the three publish modes“For backward compatibility, if streaming=True is passed, it is equivalent to setting publish_mode="CONTINUOUS".”
↩︎ publish_table and the three publish modes“you can use the SQL editor to run PostgreSQL queries against your online feature tables”
↩︎ Querying and deleting online tables“It removes the the table from both Unity Catalog and the database.”
↩︎ Querying and deleting online tables“When deployed as endpoints, these models use Unity Catalog to find appropriate features in online stores.”
↩︎ Querying and deleting online tables“Only the latest feature values for each entity ID are available in the online store for real-time applications.”
↩︎ Exam trap 1“DROP TABLE or the Python SDK command to delete a synced table do not delete the table from underlying database storage”
↩︎ Exam trap 2“Change Data Feed enabled: Required for the CONTINUOUS and TRIGGERED publish modes.”
↩︎ Exam trap 3“Only the latest feature values for each entity ID are available in the online store for real-time applications.”
↩︎ Prediction“This is an exact-match lookup on a string primary key, not a point-in-time (as-of) or temporal range query.”
↩︎ Checkpoint“Default. Incrementally updates the online table with changes from the offline table using the API or on a schedule.”
↩︎ Checkpoint - 2.
“This enables point-in-time lookups when you use create_training_set or score_batch.”
↩︎ Time series tables: the online copy keeps only the latest value