CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 4 · Lesson 43/48

    Choosing a Model Serving Approach on Databricks: Endpoints, Pipelines and Precomputed Predictions

    Identify the differences and advantages of model serving approaches: batch, realtime, and streaming

    8 min read
    2.08% of exam
    7 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Identify what real-time Model Serving endpoints provide and the trade-offs of their settings
    • Recognise how batch and streaming inference run as Databricks pipelines using AI Functions, Lakeflow pipelines and MLflow models
    • Choose between real-time serving and precomputed batch predictions when a use case needs low-latency reads

    1.Real-time serving: a model behind a REST API

    On Databricks, real-time inference runs on Model Serving. 'Each model you serve is available as a REST API that you can integrate into your web or client application.' Model Serving runs on serverless compute and 'automatically scales up or down to meet demand changes, saving infrastructure costs while optimizing latency performance.' It offers a unified REST API and the MLflow Deployment API for creating, managing and querying endpoints, plus a single UI to manage all your models and their endpoints.

    A custom model goes through three steps. First, log it in the MLflow format using a built-in flavor or pyfunc. Next, register it, preferably in Unity Catalog. Finally, create a serving endpoint for it. In the Model Serving glossary, an endpoint is a 'REST API that exposes one or more served models for inference.' Each served model is a named deployment unit with its own compute configuration.

    The main advantage of real-time serving is that it answers each request as it arrives. That speed depends on the compute being ready, which is what the next two settings control.

    Scale to zero automatically reduces resource consumption to zero when an endpoint isn't in use. It suits testing and development but not production. For production sizing, provisioned concurrency sets the maximum number of parallel requests. The glossary gives the estimate as provisioned concurrency = queries per second (QPS) × model execution time (s).

    An endpoint can also host more than one served model and split traffic between them, which lets you compare model versions online with live requests. When you update an endpoint, Model Serving keeps the existing configuration running until the new one is ready, so the update has zero downtime.

    Checkpoint 1 of 5· Check yourself

    An endpoint receives about 20 queries per second, and the model takes 0.5 seconds per prediction. Using the glossary formula, what provisioned concurrency should you estimate?

    Sources1234

    2.Batch and streaming inference as pipelines

    Batch and streaming inference don't use a per-request endpoint. They run as pipelines. In SQL, AI Functions let you 'run batch inference using task-specific AI functions or the general purpose function, ai_query.' For custom models, ai_query provides efficient batch inference for models deployed as Model Serving endpoints. Python teams can instead use custom code with Apache Spark UDFs or mlflow.pyfunc for batch inference. The sources here name these options but don't include a pandas code walkthrough.

    These pipelines can be deployed as Lakeflow pipelines, as scheduled workflows using Databricks workflows, or as 'Streaming inference workflows using Structured Streaming'. The Lakeflow example below performs incremental inference on data that is continuously updated. Its first step is a streaming table that picks up only new files from a volume.

    A Lakeflow pipeline streaming table that incrementally ingests new files before inference runssql
    CREATE OR REFRESH STREAMING TABLE news_raw
    COMMENT "Raw news articles ingested from volume."
    AS SELECT *
    FROM STREAM(read_files(
      '/Volumes/databricks_news_summarization_benchmarking_data/v01/csv',
      format => 'csv',
      header => true,
      mode => 'PERMISSIVE',
      multiLine => 'true'
    ));

    Checkpoint 2 of 5· Fill the gap

    The Python version of this pipeline ingests files incrementally. Which DataFrame reader makes it a streaming source?

    @dp.table(
      comment="Raw news articles ingested from volume."
    )
    def news_raw():
      return (
        spark. ? 
          .format("cloudFiles")

    Batch and streaming pipelines usually load the registered model from Unity Catalog by its alias rather than by a fixed version number. When the alias moves to a new version, the pipeline automatically uses that version on its next run. 'In this way the model deployment step is decoupled from inference pipelines.' Promoting a new model therefore doesn't require editing the pipeline.

    Checkpoint 3 of 5· Exam question

    A fraud-detection model must return an approve/decline decision to a checkout service within 100 milliseconds of a card being swiped, and the volume of requests varies sharply between business hours and overnight. Which approach satisfies both the latency and the variable-load requirement?

    Sources5674

    3.Choosing: when low latency doesn't require real-time serving

    Exam questions often describe a use case and ask which approach fits. A simple rule works: ask when the prediction can be computed, not only how fast it must be read. If predictions depend on inputs that exist only at request time, you need a REST endpoint. If the application needs fast lookups but the inputs are known ahead of time, there is a cheaper hybrid. Compute the predictions in batch, then 'these batch predictions can be published to an online key-value store such as DynamoDB or Cosmos DB.' The application reads from the store at low latency, and you avoid running a serving endpoint.

    Matching a use case to a serving approach
    SituationApproachWhy
    High throughput, latency tolerantBatch or streaming inferenceGenerally the most cost-effective option
    Low latency needed, but predictions can be computed offlineBatch, published to an online key-value store such as DynamoDB or Cosmos DBFast reads without a live model endpoint
    Real-time use caseModel deployed as a REST API endpoint with Model ServingEach request is scored when it arrives

    Checkpoint 4 of 5· Check yourself

    A mobile app shows each user a product recommendation within milliseconds of the screen opening. The recommendations depend only on purchase history, which is refreshed overnight. What is the most cost-effective approach that still meets the latency requirement?

    Checkpoint 5 of 5· Exam question

    A sensor-monitoring pipeline built on Lakeflow Spark Declarative Pipelines ingests IoT readings continuously via Auto Loader and needs to attach an anomaly score to each reading as it flows through the pipeline, without a person or downstream system waiting on any single call. Which serving approach fits this scenario, and what makes it a better fit than the alternatives?

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Any use case that needs low-latency predictions must use a real-time Model Serving endpoint.Why is that wrong?

      If the predictions can be computed ahead of time, batch inference can publish them to an online key-value store, and the application reads them at low latency.

      Covered in Choosing: when low latency doesn't require real-time serving

    2. 2.Turning on scale to zero is the best way to run a production real-time endpoint cheaply.Why is that wrong?

      Scale to zero is recommended for testing and development. In production, scaling up from zero adds latency and capacity is not guaranteed.

      Covered in Real-time serving: a model behind a REST API

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Each model you serve is available as a REST API that you can integrate into your web or client application.”
      ↩︎ Real-time serving: a model behind a REST API
      “The service automatically scales up or down to meet demand changes, saving infrastructure costs while optimizing latency performance.”
      ↩︎ Real-time serving: a model behind a REST API
      “Model Serving offers a unified REST API and MLflow Deployment API for CRUD and querying tasks.”
      ↩︎ Real-time serving: a model behind a REST API
    2. 2.
      “Log the model or code in the MLflow format, using either native MLflow built-in flavors or pyfunc.”
      ↩︎ Real-time serving: a model behind a REST API
    3. 3.
      “REST API that exposes one or more served models for inference.”
      ↩︎ Real-time serving: a model behind a REST API
      “Scale to zero is recommended for testing and development.”
      ↩︎ Exam trap 2
      “scale to zero is not recommended for production endpoints, as latency is greater and capacity is not guaranteed when scaled to zero.”
      ↩︎ Prediction
      “provisioned concurrency = queries per second (QPS) × model execution time (s).”
      ↩︎ Checkpoint
    4. 4.
      “You can create a single endpoint with multiple models and specify the endpoint traffic split between those models”
      ↩︎ Real-time serving: a model behind a REST API
      “Model Serving executes a zero-downtime update by keeping the existing configuration running until the new one is ready.”
      ↩︎ Real-time serving: a model behind a REST API
      “In this way the model deployment step is decoupled from inference pipelines.”
      ↩︎ Batch and streaming inference as pipelines
      “Batch or streaming inference is generally the most cost-effective option for higher throughput, higher latency use cases.”
      ↩︎ Choosing: when low latency doesn't require real-time serving
      “these batch predictions can be published to an online key-value store such as DynamoDB or Cosmos DB.”
      ↩︎ Choosing: when low latency doesn't require real-time serving
      “these batch predictions can be published to an online key-value store such as DynamoDB or Cosmos DB.”
      ↩︎ Exam trap 1
    5. 5.
      “You can run batch inference using task-specific AI functions or the general purpose function, ai_query.”
      ↩︎ Batch and streaming inference as pipelines
    6. 6.
      “ai_query provides efficient batch inference for custom models deployed as Model Serving endpoints.”
      ↩︎ Batch and streaming inference as pipelines
      “You can also use custom code with Apache Spark UDFs (example) or mlflow.pyfunc for batch inference.”
      ↩︎ Batch and streaming inference as pipelines
    7. 7.
      “The following example performs incremental batch inference using Lakeflow pipelines for when data is continuously updated.”
      ↩︎ Batch and streaming inference as pipelines

    Spotted a mistake, or was something unclear? Tell us.