What you will be able to do
- Explain the latency-versus-throughput trade-off that separates batch, streaming and real-time model serving
- Describe how batch and streaming processing differ in what data each run processes
- Compare the pros and cons of batch and streaming processing, including how each handles late-arriving data
Key concept
Latency versus throughput trade-off in model serving — Batch and streaming inference score large volumes of data cheaply but don't return predictions right away. Real-time serving puts a model behind a REST endpoint so each request gets a low-latency answer. You pick a serving approach mainly by deciding how quickly a prediction has to come back.
1.Three ways to turn a model into predictions
Once a model is trained and registered, there are three ways to use it. Batch inference scores a whole dataset in one run. Streaming inference scores data incrementally as it arrives. Real-time serving keeps the model running behind an endpoint and answers each request as it comes in.
The main thing that separates them is latency. Databricks' MLOps guidance says batch or streaming inference is generally the most cost-effective option when you need high throughput and can accept higher latency. Real-time is a different kind of deployment: 'For real-time use cases, you must set up the infrastructure to deploy the model as a REST API endpoint.' On Databricks, that infrastructure is Model Serving, which offers 'a highly available and low-latency service for deploying models'.
The approaches also differ in where the predictions end up. A batch job writes its results somewhere for later use. A real-time endpoint returns the prediction straight to the application that asked for it.
| Approach | Typical latency profile | Where predictions go |
|---|---|---|
| Batch | Higher throughput, higher latency | Tables in the production catalog, flat files, or over a JDBC connection |
| Streaming | Higher throughput, higher latency | Unity Catalog tables or message queues like Apache Kafka |
| Real-time | Low latency | Returned by a REST API endpoint managed with Model Serving |
Checkpoint 1 of 3· Check yourself
A streaming inference job has to make its predictions available to other systems. According to Databricks, where do streaming jobs typically publish predictions?
Streaming jobs write to Unity Catalog tables or message queues. Flat files and JDBC connections are the destinations listed for batch jobs, and REST responses are how real-time serving returns results.
“Streaming jobs typically publish predictions either to Unity Catalog tables or to message queues like Apache Kafka.”Source: docs.databricks.com
2.What batch and streaming actually do differently
Many people think 'streaming' means 'continuous and millisecond-fast'. Databricks defines it more broadly: streaming 'has a more expansive definition'. The engine 'can treat sources like cloud object storage and Delta Lake as streaming sources for efficient incremental processing', and streaming processing 'can be run in both triggered and continuous manners'. So a streaming job can run on a schedule and read a Delta table. What makes it streaming is the processing semantics, not how often it runs.
The real difference is state. With batch processing, the engine does not keep track of what data is already being processed in the source. Each run processes everything currently available, so in practice batch sources are partitioned, for example by day or region, to limit how much gets reprocessed. With streaming processing, 'the engine keeps track of what data is being processed and only processes new data in subsequent runs.'
Databricks illustrates this with an hourly average sales price. As a batch job, each hourly run reprocesses the earlier hours and overwrites the previous results. As a streaming job, each run processes only the rows added since the last run, and the new results are appended to the earlier ones.
In batch mode, the late hour-1 rows are processed together with the rest of the data, and the hour-1 result is overwritten with a corrected value. In streaming mode, the late rows are processed without the hour-1 data that was already handled. To fix the earlier average, the logic must have stored the sum and count from hour 1. This is why streaming gets harder for stateful operations such as joins, aggregations and deduplications.
This complexity only appears when processing is stateful. For stateless processing, such as appending new rows, late data is simply appended when it arrives. Structured Streaming, the engine behind Lakeflow pipelines, describes itself as 'a near real-time processing engine that offers end-to-end fault tolerance with exactly-once processing guarantees using familiar Spark APIs.' It lets you set trigger intervals to balance latency against cost, and it also has a real-time mode for very low-latency workloads.
Checkpoint 2 of 3· Exam question
A retail analytics team scores a churn model against the full customer table once every night and writes the predictions to a Delta table that a dashboard reads the next morning. No individual prediction needs to be available within seconds of a new customer event. Which serving approach best fits this workload, and why?
Correct answer: A — Batch inference, because loading the model with `mlflow.pyfunc.load_model` and applying it across the full table on a schedule is cheaper and simpler when low millisecond latency is not required.
- A. Nightly scoring of a full table with no sub-second latency requirement is the canonical batch inference pattern: loading the model once and applying it across the dataset (for example via a pandas UDF or `.apply`) avoids the operational overhead of a standing endpoint.
- B. A REST endpoint is built for low-latency request/response traffic; polling it for a nightly bulk job adds endpoint management and per-request overhead that a scheduled batch job does not need.
- C. Streaming inference is designed for continuously arriving records from a source like Auto Loader or a message queue; running it against a table that only changes once a day processes no new data most of the time and wastes the always-on compute.
- D. Scale-to-zero reduces idle cost but the endpoint still has to warm up and serve individual requests one at a time, which is a worse fit for scoring an entire table than a single scheduled batch pass.
3.Pros, cons and latency ranges
Databricks summarises the trade-off in a pros-and-cons table. Batch is simple and always accurate, but it is less efficient and too slow for second-level latency. Streaming is efficient and can reach milliseconds, but stateful logic gets complex, and its results are not always accurate when data arrives late or out of order.
| Semantic | Pros | Cons | Example features |
|---|---|---|---|
| Batch | Simple logic; results always accurate and reflect all available data | Reprocesses data in a batch partition; latency from hours to minutes, but not seconds or milliseconds | Materialized view; spark.read.load() and spark.write.save() |
| Streaming | Efficient, only new data is processed; latency from hours down to milliseconds | Complex stateful logic; results not always accurate with out-of-order and late-arrival data | Streaming table, append flow, sink; spark.readStream.load() and spark.writeStream.start() |
Checkpoint 3 of 3· Match them up
Match each statement to the processing semantic it describes
Tap a term, then the definition that fits it.
Batch trades efficiency for accuracy and simplicity. Streaming trades accuracy and simplicity for efficiency and lower latency.
“Slower, could handle latency requirements from hours to minutes, but not seconds or milliseconds.”Source: docs.databricks.com
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Streaming always means continuous, always-on processing from a message bus such as Kafka.Why is that wrong?
On Databricks, streaming means the engine tracks what it has already processed. It can run on a trigger rather than continuously, and it can read from Delta Lake or cloud object storage.
2.Streaming is strictly better than batch because it processes only new data.Why is that wrong?
Streaming is more efficient, but stateful logic is harder and results can be wrong when data arrives late or out of order. Batch results always reflect all available data.
Covered in Pros, cons and latency ranges
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“For real-time use cases, you must set up the infrastructure to deploy the model as a REST API endpoint.”
↩︎ Three ways to turn a model into predictions“Batch jobs typically publish predictions to tables in the production catalog, to flat files, or over a JDBC connection.”
↩︎ Three ways to turn a model into predictions“Batch or streaming inference is generally the most cost-effective option for higher throughput, higher latency use cases.”
↩︎ Key concept“Batch or streaming inference is generally the most cost-effective option for higher throughput, higher latency use cases.”
↩︎ Prediction“Streaming jobs typically publish predictions either to Unity Catalog tables or to message queues like Apache Kafka.”
↩︎ Checkpoint - 2.
“Model Serving provides a highly available and low-latency service for deploying models.”
↩︎ Three ways to turn a model into predictions - 3.
“The engine can treat sources like cloud object storage and Delta Lake as streaming sources for efficient incremental processing.”
↩︎ What batch and streaming actually do differently“With streaming processing, the engine keeps track of what data is being processed and only processes new data in subsequent runs.”
↩︎ What batch and streaming actually do differently“With batch processing, the engine does not keep track of what data is already being processed in the source.”
↩︎ What batch and streaming actually do differently“Faster, could handle latency requirements from hours to minutes, seconds, and milliseconds.”
↩︎ Pros, cons and latency ranges“Results are always accurate and reflect all the available data in the source.”
↩︎ Pros, cons and latency ranges“Streaming processing can be run in both triggered and continuous manners”
↩︎ Exam trap 1“Results cannot always be accurate, considering out-of-order and late-arrival data.”
↩︎ Exam trap 2“Streaming processing can be run in both triggered and continuous manners”
↩︎ Prediction“Slower, could handle latency requirements from hours to minutes, but not seconds or milliseconds.”
↩︎ Checkpoint - 4.
“Apache Spark Structured Streaming is a near real-time processing engine that offers end-to-end fault tolerance with exactly-once processing guarantees using familiar Spark APIs.”
↩︎ What batch and streaming actually do differently“Set trigger intervals to balance latency and cost for your processing requirements.”
↩︎ What batch and streaming actually do differently