CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 4 · Lesson 43/48

    Batch vs Streaming vs Real-Time Inference: Latency, Cost and Semantics

    Identify the differences and advantages of model serving approaches: batch, realtime, and streaming

    7 min read
    2.08% of exam
    4 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain the latency-versus-throughput trade-off that separates batch, streaming and real-time model serving
    • Describe how batch and streaming processing differ in what data each run processes
    • Compare the pros and cons of batch and streaming processing, including how each handles late-arriving data

    Key concept

    Latency versus throughput trade-off in model serving — Batch and streaming inference score large volumes of data cheaply but don't return predictions right away. Real-time serving puts a model behind a REST endpoint so each request gets a low-latency answer. You pick a serving approach mainly by deciding how quickly a prediction has to come back.

    1.Three ways to turn a model into predictions

    Once a model is trained and registered, there are three ways to use it. Batch inference scores a whole dataset in one run. Streaming inference scores data incrementally as it arrives. Real-time serving keeps the model running behind an endpoint and answers each request as it comes in.

    The main thing that separates them is latency. Databricks' MLOps guidance says batch or streaming inference is generally the most cost-effective option when you need high throughput and can accept higher latency. Real-time is a different kind of deployment: 'For real-time use cases, you must set up the infrastructure to deploy the model as a REST API endpoint.' On Databricks, that infrastructure is Model Serving, which offers 'a highly available and low-latency service for deploying models'.

    The approaches also differ in where the predictions end up. A batch job writes its results somewhere for later use. A real-time endpoint returns the prediction straight to the application that asked for it.

    Where each serving approach typically sends its predictions
    ApproachTypical latency profileWhere predictions go
    BatchHigher throughput, higher latencyTables in the production catalog, flat files, or over a JDBC connection
    StreamingHigher throughput, higher latencyUnity Catalog tables or message queues like Apache Kafka
    Real-timeLow latencyReturned by a REST API endpoint managed with Model Serving

    Checkpoint 1 of 3· Check yourself

    A streaming inference job has to make its predictions available to other systems. According to Databricks, where do streaming jobs typically publish predictions?

    Sources12

    2.What batch and streaming actually do differently

    Many people think 'streaming' means 'continuous and millisecond-fast'. Databricks defines it more broadly: streaming 'has a more expansive definition'. The engine 'can treat sources like cloud object storage and Delta Lake as streaming sources for efficient incremental processing', and streaming processing 'can be run in both triggered and continuous manners'. So a streaming job can run on a schedule and read a Delta table. What makes it streaming is the processing semantics, not how often it runs.

    The real difference is state. With batch processing, the engine does not keep track of what data is already being processed in the source. Each run processes everything currently available, so in practice batch sources are partitioned, for example by day or region, to limit how much gets reprocessed. With streaming processing, 'the engine keeps track of what data is being processed and only processes new data in subsequent runs.'

    Databricks illustrates this with an hourly average sales price. As a batch job, each hourly run reprocesses the earlier hours and overwrites the previous results. As a streaming job, each run processes only the rows added since the last run, and the new results are appended to the earlier ones.

    This complexity only appears when processing is stateful. For stateless processing, such as appending new rows, late data is simply appended when it arrives. Structured Streaming, the engine behind Lakeflow pipelines, describes itself as 'a near real-time processing engine that offers end-to-end fault tolerance with exactly-once processing guarantees using familiar Spark APIs.' It lets you set trigger intervals to balance latency against cost, and it also has a real-time mode for very low-latency workloads.

    Checkpoint 2 of 3· Exam question

    A retail analytics team scores a churn model against the full customer table once every night and writes the predictions to a Delta table that a dashboard reads the next morning. No individual prediction needs to be available within seconds of a new customer event. Which serving approach best fits this workload, and why?

    Sources34

    3.Pros, cons and latency ranges

    Databricks summarises the trade-off in a pros-and-cons table. Batch is simple and always accurate, but it is less efficient and too slow for second-level latency. Streaming is efficient and can reach milliseconds, but stateful logic gets complex, and its results are not always accurate when data arrives late or out of order.

    Batch vs streaming processing semantics in Databricks
    SemanticProsConsExample features
    BatchSimple logic; results always accurate and reflect all available dataReprocesses data in a batch partition; latency from hours to minutes, but not seconds or millisecondsMaterialized view; spark.read.load() and spark.write.save()
    StreamingEfficient, only new data is processed; latency from hours down to millisecondsComplex stateful logic; results not always accurate with out-of-order and late-arrival dataStreaming table, append flow, sink; spark.readStream.load() and spark.writeStream.start()

    Checkpoint 3 of 3· Match them up

    Match each statement to the processing semantic it describes

    Tap a term, then the definition that fits it.

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Streaming always means continuous, always-on processing from a message bus such as Kafka.Why is that wrong?

      On Databricks, streaming means the engine tracks what it has already processed. It can run on a trigger rather than continuously, and it can read from Delta Lake or cloud object storage.

      Covered in What batch and streaming actually do differently

    2. 2.Streaming is strictly better than batch because it processes only new data.Why is that wrong?

      Streaming is more efficient, but stateful logic is harder and results can be wrong when data arrives late or out of order. Batch results always reflect all available data.

      Covered in Pros, cons and latency ranges

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “For real-time use cases, you must set up the infrastructure to deploy the model as a REST API endpoint.”
      ↩︎ Three ways to turn a model into predictions
      “Batch jobs typically publish predictions to tables in the production catalog, to flat files, or over a JDBC connection.”
      ↩︎ Three ways to turn a model into predictions
      “Batch or streaming inference is generally the most cost-effective option for higher throughput, higher latency use cases.”
      ↩︎ Key concept
      “Batch or streaming inference is generally the most cost-effective option for higher throughput, higher latency use cases.”
      ↩︎ Prediction
      “Streaming jobs typically publish predictions either to Unity Catalog tables or to message queues like Apache Kafka.”
      ↩︎ Checkpoint
    2. 2.
      “Model Serving provides a highly available and low-latency service for deploying models.”
      ↩︎ Three ways to turn a model into predictions
    3. 3.
      “The engine can treat sources like cloud object storage and Delta Lake as streaming sources for efficient incremental processing.”
      ↩︎ What batch and streaming actually do differently
      “With streaming processing, the engine keeps track of what data is being processed and only processes new data in subsequent runs.”
      ↩︎ What batch and streaming actually do differently
      “With batch processing, the engine does not keep track of what data is already being processed in the source.”
      ↩︎ What batch and streaming actually do differently
      “Faster, could handle latency requirements from hours to minutes, seconds, and milliseconds.”
      ↩︎ Pros, cons and latency ranges
      “Results are always accurate and reflect all the available data in the source.”
      ↩︎ Pros, cons and latency ranges
      “Streaming processing can be run in both triggered and continuous manners”
      ↩︎ Exam trap 1
      “Results cannot always be accurate, considering out-of-order and late-arrival data.”
      ↩︎ Exam trap 2
      “Streaming processing can be run in both triggered and continuous manners”
      ↩︎ Prediction
      “Slower, could handle latency requirements from hours to minutes, but not seconds or milliseconds.”
      ↩︎ Checkpoint
    4. 4.
      “Apache Spark Structured Streaming is a near real-time processing engine that offers end-to-end fault tolerance with exactly-once processing guarantees using familiar Spark APIs.”
      ↩︎ What batch and streaming actually do differently
      “Set trigger intervals to balance latency and cost for your processing requirements.”
      ↩︎ What batch and streaming actually do differently

    Continue to page 2 of 2

    Choosing a Model Serving Approach on Databricks: Endpoints, Pipelines and Precomputed Predictions

    Spotted a mistake, or was something unclear? Tell us.