What you will be able to do
- Identify what real-time Model Serving endpoints provide and the trade-offs of their settings
- Recognise how batch and streaming inference run as Databricks pipelines using AI Functions, Lakeflow pipelines and MLflow models
- Choose between real-time serving and precomputed batch predictions when a use case needs low-latency reads
1.Real-time serving: a model behind a REST API
On Databricks, real-time inference runs on Model Serving. 'Each model you serve is available as a REST API that you can integrate into your web or client application.' Model Serving runs on serverless compute and 'automatically scales up or down to meet demand changes, saving infrastructure costs while optimizing latency performance.' It offers a unified REST API and the MLflow Deployment API for creating, managing and querying endpoints, plus a single UI to manage all your models and their endpoints.
A custom model goes through three steps. First, log it in the MLflow format using a built-in flavor or pyfunc. Next, register it, preferably in Unity Catalog. Finally, create a serving endpoint for it. In the Model Serving glossary, an endpoint is a 'REST API that exposes one or more served models for inference.' Each served model is a named deployment unit with its own compute configuration.
The main advantage of real-time serving is that it answers each request as it arrives. That speed depends on the compute being ready, which is what the next two settings control.
Scale to zero automatically reduces resource consumption to zero when an endpoint isn't in use. It suits testing and development but not production. For production sizing, provisioned concurrency sets the maximum number of parallel requests. The glossary gives the estimate as provisioned concurrency = queries per second (QPS) × model execution time (s).
An endpoint can also host more than one served model and split traffic between them, which lets you compare model versions online with live requests. When you update an endpoint, Model Serving keeps the existing configuration running until the new one is ready, so the update has zero downtime.
Checkpoint 1 of 5· Check yourself
An endpoint receives about 20 queries per second, and the model takes 0.5 seconds per prediction. Using the glossary formula, what provisioned concurrency should you estimate?
Provisioned concurrency is QPS multiplied by model execution time in seconds: 20 × 0.5 = 10.
“provisioned concurrency = queries per second (QPS) × model execution time (s).”Source: docs.databricks.com
2.Batch and streaming inference as pipelines
Batch and streaming inference don't use a per-request endpoint. They run as pipelines. In SQL, AI Functions let you 'run batch inference using task-specific AI functions or the general purpose function, ai_query.' For custom models, ai_query provides efficient batch inference for models deployed as Model Serving endpoints. Python teams can instead use custom code with Apache Spark UDFs or mlflow.pyfunc for batch inference. The sources here name these options but don't include a pandas code walkthrough.
These pipelines can be deployed as Lakeflow pipelines, as scheduled workflows using Databricks workflows, or as 'Streaming inference workflows using Structured Streaming'. The Lakeflow example below performs incremental inference on data that is continuously updated. Its first step is a streaming table that picks up only new files from a volume.
CREATE OR REFRESH STREAMING TABLE news_raw
COMMENT "Raw news articles ingested from volume."
AS SELECT *
FROM STREAM(read_files(
'/Volumes/databricks_news_summarization_benchmarking_data/v01/csv',
format => 'csv',
header => true,
mode => 'PERMISSIVE',
multiLine => 'true'
));Checkpoint 2 of 5· Fill the gap
The Python version of this pipeline ingests files incrementally. Which DataFrame reader makes it a streaming source?
@dp.table(
comment="Raw news articles ingested from volume."
)
def news_raw():
return (
spark. ?
.format("cloudFiles")spark.readStream with the cloudFiles format (Auto Loader) reads only files that have arrived since the last run. spark.read would reprocess everything on every run, which is batch behaviour.
Source: docs.databricks.comBatch and streaming pipelines usually load the registered model from Unity Catalog by its alias rather than by a fixed version number. When the alias moves to a new version, the pipeline automatically uses that version on its next run. 'In this way the model deployment step is decoupled from inference pipelines.' Promoting a new model therefore doesn't require editing the pipeline.
Checkpoint 3 of 5· Exam question
A fraud-detection model must return an approve/decline decision to a checkout service within 100 milliseconds of a card being swiped, and the volume of requests varies sharply between business hours and overnight. Which approach satisfies both the latency and the variable-load requirement?
Correct answer: A — A Databricks Model Serving REST endpoint, because it exposes the model over a low-latency API and automatically scales compute up or down as request volume changes.
- A. Model Serving endpoints are purpose-built for synchronous, low-latency request/response traffic and scale automatically with demand, which matches both the 100ms requirement and the swings between business hours and overnight.
- B. Batch scoring the previous day's transactions produces decisions for events that already happened, which cannot support a checkout flow that needs a decision at the moment of the swipe.
- C. A streaming pipeline writing to a Delta table adds the latency of a micro-batch or continuous trigger and a table read before the checkout service sees a result, which is not a synchronous request/response path suited to a 100ms budget.
- D. A fixed-size always-on cluster avoids cold starts but does not scale down during overnight lulls or scale up during peak hours, so it either overpays for idle capacity or falls behind under load.
3.Choosing: when low latency doesn't require real-time serving
Exam questions often describe a use case and ask which approach fits. A simple rule works: ask when the prediction can be computed, not only how fast it must be read. If predictions depend on inputs that exist only at request time, you need a REST endpoint. If the application needs fast lookups but the inputs are known ahead of time, there is a cheaper hybrid. Compute the predictions in batch, then 'these batch predictions can be published to an online key-value store such as DynamoDB or Cosmos DB.' The application reads from the store at low latency, and you avoid running a serving endpoint.
| Situation | Approach | Why |
|---|---|---|
| High throughput, latency tolerant | Batch or streaming inference | Generally the most cost-effective option |
| Low latency needed, but predictions can be computed offline | Batch, published to an online key-value store such as DynamoDB or Cosmos DB | Fast reads without a live model endpoint |
| Real-time use case | Model deployed as a REST API endpoint with Model Serving | Each request is scored when it arrives |
Checkpoint 4 of 5· Check yourself
A mobile app shows each user a product recommendation within milliseconds of the screen opening. The recommendations depend only on purchase history, which is refreshed overnight. What is the most cost-effective approach that still meets the latency requirement?
The predictions can be computed offline, so batch scoring followed by publishing to an online key-value store gives low-latency reads without the cost of an always-on endpoint.
“these batch predictions can be published to an online key-value store such as DynamoDB or Cosmos DB.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A sensor-monitoring pipeline built on Lakeflow Spark Declarative Pipelines ingests IoT readings continuously via Auto Loader and needs to attach an anomaly score to each reading as it flows through the pipeline, without a person or downstream system waiting on any single call. Which serving approach fits this scenario, and what makes it a better fit than the alternatives?
Correct answer: A — Streaming inference, because applying a registered MLflow model as a transformation inside the declarative pipeline scores each new record as it arrives, with no separate scheduled job needed.
- A. Scoring records as they flow through an already-continuous pipeline by loading the model as a UDF inside the pipeline transformation is the streaming inference pattern, matching the continuous ingestion and avoiding a separate synchronous call per row.
- B. Scheduling batch scoring on top of a continuously arriving stream introduces artificial delay between when a reading lands and when it is scored, undoing the low-latency benefit that a continuously running pipeline already provides.
- C. Calling a REST endpoint synchronously for every row inside a high-throughput streaming pipeline adds network round-trip overhead per record and turns the pipeline's throughput into a function of endpoint request latency.
- D. Splitting continuously arriving sensor data between a batch path and a real-time path adds two serving mechanisms to maintain for one continuous data source when the pipeline can score every record inline as it passes through.
Sources4
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Any use case that needs low-latency predictions must use a real-time Model Serving endpoint.Why is that wrong?
If the predictions can be computed ahead of time, batch inference can publish them to an online key-value store, and the application reads them at low latency.
Covered in Choosing: when low latency doesn't require real-time serving
2.Turning on scale to zero is the best way to run a production real-time endpoint cheaply.Why is that wrong?
Scale to zero is recommended for testing and development. In production, scaling up from zero adds latency and capacity is not guaranteed.
Covered in Real-time serving: a model behind a REST API
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Each model you serve is available as a REST API that you can integrate into your web or client application.”
↩︎ Real-time serving: a model behind a REST API“The service automatically scales up or down to meet demand changes, saving infrastructure costs while optimizing latency performance.”
↩︎ Real-time serving: a model behind a REST API“Model Serving offers a unified REST API and MLflow Deployment API for CRUD and querying tasks.”
↩︎ Real-time serving: a model behind a REST API - 2.
“Log the model or code in the MLflow format, using either native MLflow built-in flavors or pyfunc.”
↩︎ Real-time serving: a model behind a REST API - 3.
“REST API that exposes one or more served models for inference.”
↩︎ Real-time serving: a model behind a REST API“Scale to zero is recommended for testing and development.”
↩︎ Exam trap 2“scale to zero is not recommended for production endpoints, as latency is greater and capacity is not guaranteed when scaled to zero.”
↩︎ Prediction“provisioned concurrency = queries per second (QPS) × model execution time (s).”
↩︎ Checkpoint - 4.
“You can create a single endpoint with multiple models and specify the endpoint traffic split between those models”
↩︎ Real-time serving: a model behind a REST API“Model Serving executes a zero-downtime update by keeping the existing configuration running until the new one is ready.”
↩︎ Real-time serving: a model behind a REST API“In this way the model deployment step is decoupled from inference pipelines.”
↩︎ Batch and streaming inference as pipelines“Batch or streaming inference is generally the most cost-effective option for higher throughput, higher latency use cases.”
↩︎ Choosing: when low latency doesn't require real-time serving“these batch predictions can be published to an online key-value store such as DynamoDB or Cosmos DB.”
↩︎ Choosing: when low latency doesn't require real-time serving“these batch predictions can be published to an online key-value store such as DynamoDB or Cosmos DB.”
↩︎ Exam trap 1 - 5.
“You can run batch inference using task-specific AI functions or the general purpose function, ai_query.”
↩︎ Batch and streaming inference as pipelines - 6.
“ai_query provides efficient batch inference for custom models deployed as Model Serving endpoints.”
↩︎ Batch and streaming inference as pipelines“You can also use custom code with Apache Spark UDFs (example) or mlflow.pyfunc for batch inference.”
↩︎ Batch and streaming inference as pipelines - 7.
“The following example performs incremental batch inference using Lakeflow pipelines for when data is continuously updated.”
↩︎ Batch and streaming inference as pipelines