CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 1 · Lesson 10/48

    Offline vs Online Feature Tables in Databricks

    Describe the differences between online and offline feature tables

    9 min read
    2.08% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Describe what the offline feature store holds and which workloads read from it
    • Describe what an online feature table is, what backs it and why it exists
    • Identify which store batch inference and real-time inference each read from
    • Recognise which feature tables cannot be published online and where third-party online stores fit

    Key concept

    Offline vs online feature store — The offline store is your Delta feature tables in Unity Catalog, used for discovery, training and batch scoring. The online store is a low-latency copy of those tables, published for real-time serving. Which store gets read depends on whether inference is batch or real-time.

    1.The offline store: Delta tables for discovery, training and batch scoring

    Every feature in Databricks Feature Store starts life in the offline store. With Unity Catalog enabled, any Delta table that has a primary key constraint can act as a feature table. You can build it with Databricks SQL, the Python FeatureEngineeringClient or Lakeflow pipelines, and you address it with the usual three-level name <catalog-name>.<schema-name>.<table-name>. The glossary sets out what this store is for: the offline store "is used for feature discovery, model training, and batch inference. It contains feature tables materialized as Delta tables."

    Each of those jobs works in bulk. Discovery means browsing and governing tables in Unity Catalog. Training joins whole feature tables to a label DataFrame. Batch inference scores large sets of new rows in one pass. None of them needs one row back in milliseconds, so ordinary Delta storage handles them well. Your own pipelines keep the offline table current, with batch writes or Structured Streaming, and Databricks recommends a scheduled job that refreshes feature tables regularly, for example once a day.

    Checkpoint 1 of 6· Check yourself

    Which set of workloads is the offline feature store used for?

    Sources12

    2.The online store: a low-latency published copy

    The Databricks Online Feature Store is a high-performance, scalable store for serving feature data to online applications and real-time models. It runs on Databricks Lakebase, and it "provides low-latency access to feature data at a high scale while maintaining governance, lineage, and consistency with your offline feature tables." You don't write features into it directly. You *publish* Unity Catalog tables into it, either as a sync or by streaming feature tables from the offline store to the online store. The online table is therefore a copy of an offline table, kept in sync with it, rather than a second place where features are defined.

    Offline vs online feature tables at a glance
    AspectOffline feature tablesOnline feature tables
    StorageFeature tables materialized as Delta tables in Unity CatalogDatabricks Online Feature Store powered by Lakebase (third-party online stores are also supported)
    Used forFeature discovery, model training, batch inferenceServing features to online applications and real-time ML models
    How data gets inYour feature pipelines write to them (batch or streaming)Published (or streamed) from offline feature tables
    Read at inferenceBatch inference: values joined with new data before scoringReal-time inference: values retrieved from the online store
    GovernanceUnity Catalog Delta tables with a primary keyUnity Catalog entities that track lineage to the source tables

    Checkpoint 2 of 6· Match them up

    Match each workload to the feature store that serves it

    Tap a term, then the definition that fits it.

    Sources1

    3.Which store each inference path reads

    When you log a model with FeatureEngineeringClient.log_model, the model keeps references to the features it was trained on. At inference the caller sends only the primary key, for example user_id, and the model fetches the feature values it needs. Where those values come from depends on how you score. "In batch inference, feature values are retrieved from the offline store and joined with new data prior to scoring. In real-time inference, feature values are retrieved from the online store."

    On the real-time path, a model serving endpoint "automatically uses the entity IDs in the request data to look up pre-computed features from the online store". It finds the right online tables through Unity Catalog, which resolves lineage from the served model back to its training features. The docs describe this as automatic with no setup required. When a scoring request arrives, Model Serving retrieves the published feature values, so predictions always use the most recent ones. A model trained on offline tables can therefore serve in real time without code changes, provided its features have been published.

    Checkpoint 3 of 6· Put it in order

    Put the typical Feature Store workflow in order, from raw features to real-time scoring

    1. 1.Publish the features to an online feature store
    2. 2.Register the model in Model Registry
    3. 3.Train and log a model using the feature table
    4. 4.The serving endpoint looks up pre-computed features from the online store using entity IDs in the request
    5. 5.Create a Delta table in Unity Catalog that has a primary key

    Checkpoint 4 of 6· Exam question

    A fraud-detection model is deployed to a real-time Model Serving endpoint that must score incoming transactions within tens of milliseconds. The engineering team has an existing offline feature table in Unity Catalog that holds the customer-level features the model needs, but the endpoint currently cannot retrieve those features fast enough at inference time. What should the team do to give the endpoint low-latency access to these features?

    Sources13

    4.What can and cannot be served online

    Not every offline feature table can have an online counterpart. Feature tables based on views "can be used for offline model training and evaluation. They cannot be published to online stores." Because they can't be published, features from those tables, and models built on them, cannot be served.

    FeatureSpecs follow the same split. A FeatureSpec bundles FeatureLookups and FeatureFunctions into one unit that you can use in training or deploy behind a Feature Serving endpoint. However, "A FeatureSpec always references the offline feature tables, but they must be published to an online store for real-time serving scenarios." Both ways of authoring features, Feature Views and feature tables, produce Unity Catalog-governed features that can be published to the Online Feature Store.

    The online store doesn't have to be Databricks'. Feature tables can also be published to third-party stores: Amazon DynamoDB, Amazon Aurora (MySQL-compatible) and Amazon RDS MySQL. Even so, Databricks recommends its own Online Feature Stores for real-time serving. With a third-party store the online copy can be laid out differently from the offline table. For example, the DynamoDB online store keeps primary keys as one combined key in the column _feature_store_internal__primary_keys.

    Checkpoint 5 of 6· Check yourself

    A team defines a feature table on top of a view and wants to serve a model trained on it from a real-time endpoint. What happens?

    Checkpoint 6 of 6· Exam question

    A data scientist is assembling a training set for a churn model using several feature tables joined to historical label events, and each label row must be matched to feature values as they existed at that row's timestamp to avoid leaking future information. Which approach correctly builds this training set?

    Sources145

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A model logged with feature metadata reads from the online store for every kind of inference, including batch scoring.Why is that wrong?

      Batch inference reads feature values from the offline store and joins them with the new data. Only real-time inference reads from the online store.

      Covered in Which store each inference path reads

    2. 2.Any feature table that works for training can also be published online and served.Why is that wrong?

      Feature tables based on views are limited to offline training and evaluation. They cannot be published to online stores, so their features cannot be served.

      Covered in What can and cannot be served online

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “The offline feature store is used for feature discovery, model training, and batch inference. It contains feature tables materialized as Delta tables.”
      ↩︎ The offline store: Delta tables for discovery, training and batch scoring
      “it provides low-latency access to feature data at a high scale while maintaining governance, lineage, and consistency with your offline feature tables”
      ↩︎ The online store: a low-latency published copy
      “You can also stream feature tables from the offline store to an online store.”
      ↩︎ The online store: a low-latency published copy
      “automatically uses the entity IDs in the request data to look up pre-computed features from the online store”
      ↩︎ Which store each inference path reads
      “A FeatureSpec always references the offline feature tables, but they must be published to an online store for real-time serving scenarios.”
      ↩︎ What can and cannot be served online
      “In batch inference, feature values are retrieved from the offline store and joined with new data prior to scoring.”
      ↩︎ Key concept
      “In real-time inference, feature values are retrieved from the online store.”
      ↩︎ Exam trap 1
      “The offline feature store is used for feature discovery, model training, and batch inference.”
      ↩︎ Checkpoint
      “These tables are also Unity Catalog entities that natively track lineage to the source tables.”
      ↩︎ Prediction
      “The offline feature store is used for feature discovery, model training, and batch inference.”
      ↩︎ Checkpoint
      “For real-time serving use cases, publish the features to an online feature store.”
      ↩︎ Checkpoint
    2. 2.
      “In Unity Catalog, any Delta table with a primary key constraint can serve as a feature table.”
      ↩︎ The offline store: Delta tables for discovery, training and batch scoring
      “Features from these tables and models based on these features cannot be served.”
      ↩︎ Exam trap 2
      “Feature tables based on views can be used for offline model training and evaluation. They cannot be published to online stores.”
      ↩︎ Checkpoint
    3. 3.
      “When a scoring request comes in to the model, Model Serving automatically retrieves the published feature values needed by the model.”
      ↩︎ Which store each inference path reads
    4. 4.
      “For real-time serving of feature values, Databricks recommends using Databricks Online Feature Stores.”
      ↩︎ What can and cannot be served online
    5. 5.
      “The DynamoDB online store uses a different schema than the offline store.”
      ↩︎ What can and cannot be served online

    Continue to page 2 of 2

    Publishing Offline Feature Tables to a Databricks Online Store

    Spotted a mistake, or was something unclear? Tell us.