What you will be able to do
- Recognise Delta Live Tables (DLT) under its current name, Lakeflow pipelines, and map the old dlt Python API to the dp API
- Choose between a streaming table, a materialized view and a view for the output of a scoring step
- Explain how triggered and continuous pipeline modes affect how fresh streaming predictions are, and what they cost
Key concept
Streaming table — A streaming table is a Delta table that a pipeline adds to incrementally. Each new row from an append-only source is processed once, which suits scoring a growing stream of records. Rows that were already written are not processed again.
1.Delta Live Tables is now Lakeflow pipelines
The exam objective still says "Delta Live Tables", but the documentation uses a newer name. Databricks has renamed the product, and existing code does not need to be migrated. DLT code keeps working as it is. The new names matter because current examples, including the ones that apply MLflow models, use the new Python module and decorators.
Switching to dp changes the decorators too. @dlt becomes @dp, @view becomes @temporary_view, and table creation is now split into two decorators. @table creates streaming tables and the new @materialized_view creates materialized views. Some traces of the old name remain: the classic SKUs still start with DLT, event log schemas with dlt in the name have not changed, and the old Python APIs still run, although Databricks recommends the new names.
| DLT (legacy) | Lakeflow pipelines | What it creates |
|---|---|---|
| import dlt | from pyspark import pipelines as dp | The pipelines Python module |
| @dlt | @dp | The decorator prefix |
| @table | @table | A streaming table |
| (no separate decorator) | @dp.materialized_view | A materialized view |
| @view | @temporary_view | A view that is not persisted |
Checkpoint 1 of 4· Check yourself
A colleague says their old DLT notebooks must be rewritten before they can run as Lakeflow pipelines. What do the docs say?
The rename does not break anything. DLT code keeps running, and Databricks recommends the new API names for future compatibility.
“If you have previously used DLT, there is no migration required to use Lakeflow pipelines: your code will still work.”Source: docs.databricks.com
Sources1
2.Where streaming predictions land: streaming tables vs. materialized views
A pipeline writes to three kinds of dataset, and each one processes records differently. That choice decides whether your inference step scores only new records or recomputes everything. A streaming table processes each record once from an append-only source. A materialized view is recomputed as needed so that it reflects the current state of the data. A view is evaluated on demand and not persisted. All tables a pipeline creates and manages are Delta tables, with Delta Lake guarantees such as ACID transactions and time travel.
For streaming inference, the streaming table is the natural fit. The docs recommend one when a query reads from a data source that keeps growing, when results should be computed incrementally, and when the pipeline needs high throughput and low latency. On each update, the flows that feed a streaming table read only the changed information in the streaming source and append the new results. A streaming table must read from a streaming source, such as files loaded with Auto Loader or a message bus like Apache Kafka, Azure Event Hubs or Google Pub/Sub.
@dp.table
def streaming_bronze():
return (
# Since this is a streaming source, this table is incremental.
spark.readStream.format("cloudFiles")
.option("cloudFiles.format", "json")
.load("s3://path/to/raw/data")
)
@dp.table
def streaming_silver():
# Since we read the bronze table as a stream, this silver table is also
# updated incrementally.
return spark.readStream.table("streaming_bronze").where(...)
@dp.materialized_view
def live_gold():
# This table will be recomputed completely by reading the whole silver table
# when it is updated.
return spark.read.table("streaming_silver").groupBy("user_id").count()Streaming tables have one condition attached. They are built for streams over bounded state, meaning streams that end naturally or that are bounded by a watermark. An unbounded stream with no watermark can make the pipeline fail from memory pressure. If the source data is updated or deleted rather than only growing, or the query performs joins or aggregations that must stay correct, the docs point you to a materialized view instead.
Checkpoint 2 of 4· Exam question
A machine learning engineer is building a Lakeflow Spark Declarative Pipeline (the rebranded Delta Live Tables) that must score newly arriving rows in the `bronze_orders` table with a registered MLflow model as soon as they land, without recomputing predictions for orders that were already scored on a previous run. Which approach correctly implements this streaming inference pattern?
Correct answer: A — Decorate the scoring function with `@dp.table`, read the source using `spark.readStream.table("bronze_orders")`, load the model with `mlflow.pyfunc.spark_udf`, and append predictions with `withColumn`.
- A. This is correct: `@dp.table` backed by a streaming read processes only the new micro-batch of rows on each run, and loading the model once with `mlflow.pyfunc.spark_udf` lets it be applied as an ordinary column transformation on that incremental data.
- B. A materialized view is recomputed from its complete upstream input each time the pipeline runs, so every order would be rescored on every run instead of only the new ones, defeating the incremental goal.
- C. Calling `.toPandas()` forces a full collect of the streaming data to the driver, which is unsupported inside a dataset definition and breaks the pipeline's distributed, incremental execution model.
- D. This is unnecessary and inaccurate: dataset functions can call `mlflow.pyfunc.spark_udf` directly, so routing scoring through an external `foreachBatch` sink outside the pipeline is not required to achieve incremental inference.
Checkpoint 3 of 4· Check yourself
New sensor readings land continuously in cloud storage, and you want each reading scored once, as soon as possible after it arrives. Which pipeline dataset fits best?
The source only grows, each record should be processed once, and latency matters. The docs list exactly these as the reasons to use a streaming table.
“A query is defined against a data source that is continuously or incrementally growing.”Source: docs.databricks.com
3.How often predictions refresh: triggered vs. continuous mode
Running a streaming table does not by itself make the pipeline run all the time. A pipeline runs in either triggered or continuous mode. The mode controls when new predictions appear. In triggered mode, the pipeline processes the data that is available when the update starts and then stops. In continuous mode, it keeps processing new data as it arrives. To avoid wasted work, a continuous pipeline monitors the Delta tables it depends on and runs an update only when their contents have changed.
| Question | Triggered | Continuous |
|---|---|---|
| When does the update stop? | Automatically when complete | Runs until manually stopped |
| What data is processed? | Data available when the update starts | All data as it arrives at configured sources |
| Freshness it suits | Updates every 10 minutes, hourly, or daily | Updates every 10 seconds to a few minutes |
| Cost profile | Cluster runs only long enough to update | Always-running cluster: more expensive, lower latency |
How you get continuous execution matters too. Databricks recommends wrapping the pipeline in a continuous job instead of setting the pipeline's own Pipeline mode to continuous. The job controls the pipeline's lifecycle and unlocks serverless performance modes, such as Standard mode, that the built-in continuous setting does not support. A job's execution mode also overrides the pipeline's setting, so the docs advise leaving Pipeline mode at the default, triggered, when a continuous job wraps the pipeline. In continuous runs, pipelines.trigger.interval controls how often each flow starts an update. Standalone streaming tables, created outside a Lakeflow pipeline, always refresh in triggered mode. For sub-second, end-to-end latency, the docs also describe a separate real-time mode.
Checkpoint 4 of 4· Match them up
Match each setting to its behaviour
Tap a term, then the definition that fits it.
Triggered and continuous modes differ in when updates stop and which data they process. A continuous job is the recommended way to run continuously, and the trigger interval applies only to continuous runs.
“Databricks recommends running continuous pipelines with a continuous job rather than setting the value of Pipeline mode to continuous.”Source: docs.databricks.com
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.In the dp API, @dp.table creates a generic table that may be batch or streaming, the same way @dlt.table did.Why is that wrong?
In Lakeflow pipelines, @table creates streaming tables. Materialized views now have their own decorator, @materialized_view.
Covered in Delta Live Tables is now Lakeflow pipelines
2.Streaming inference with a streaming table requires the pipeline to run in continuous mode.Why is that wrong?
Pipeline mode and dataset type are independent. A streaming table can be refreshed by a triggered pipeline, which processes only the data available when the update starts.
Covered in How often predictions refresh: triggered vs. continuous mode
3.The recommended way to run a pipeline continuously is to set its Pipeline mode to continuous.Why is that wrong?
Databricks recommends a continuous job instead. The job controls the execution mode and overrides the pipeline's own setting.
Covered in How often predictions refresh: triggered vs. continuous mode
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The product formerly known as Delta Live Tables (DLT) has been updated to Lakeflow pipelines.”
↩︎ Delta Live Tables is now Lakeflow pipelines“The @table decorator is now used to create streaming tables, and the new @materialized_view decorator is used to create materialized views.”
↩︎ Delta Live Tables is now Lakeflow pipelines“The @table decorator is now used to create streaming tables, and the new @materialized_view decorator is used to create materialized views.”
↩︎ Exam trap 1“In Python code, references to import dlt can be replaced with from pyspark import pipelines as dp”
↩︎ Prediction“If you have previously used DLT, there is no migration required to use Lakeflow pipelines: your code will still work.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/ldp/conceptsOfficial docs
“Each record is processed exactly one time, assuming an append-only source.”
↩︎ Where streaming predictions land: streaming tables vs. materialized views“Streaming tables are always defined against streaming sources.”
↩︎ Where streaming predictions land: streaming tables vs. materialized views“All tables created and managed by pipelines are Delta tables.”
↩︎ Where streaming predictions land: streaming tables vs. materialized views“Each record is processed exactly one time, assuming an append-only source.”
↩︎ Key concept“A query is defined against a data source that is continuously or incrementally growing.”
↩︎ Checkpoint - 3.
“the flows associated with a streaming table read the changed information in a streaming source, and append new information to that table”
↩︎ Where streaming predictions land: streaming tables vs. materialized views“An unbounded stream that does not have a watermark can cause a pipeline to fail due to memory pressure.”
↩︎ Where streaming predictions land: streaming tables vs. materialized views - 4.
“If the pipeline uses triggered mode, the system stops after refreshing all tables based on the data available when the update started.”
↩︎ How often predictions refresh: triggered vs. continuous mode“Continuous pipelines require an always-running cluster, which is more expensive but reduces processing latency.”
↩︎ How often predictions refresh: triggered vs. continuous mode“Both materialized views and streaming tables can be updated in either pipeline mode.”
↩︎ Exam trap 2“Databricks recommends running continuous pipelines with a continuous job rather than setting the value of Pipeline mode to continuous.”
↩︎ Exam trap 3“Pipeline mode is independent of the type of table being computed.”
↩︎ Prediction“Databricks recommends running continuous pipelines with a continuous job rather than setting the value of Pipeline mode to continuous.”
↩︎ Checkpoint