CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 4 · Lesson 44/48

    Sizing, Scaling and Updating Custom Model Serving Endpoints

    Deploy a custom model to a model endpoint

    10 min read
    2.08% of exam
    2 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Choose a CPU or GPU workload type for a custom model
    • Size provisioned concurrency from QPS and model execution time
    • Decide when scale to zero is appropriate
    • Update a live endpoint to a new model version without downtime

    1.Choose a workload type: CPU, larger CPU or GPU

    Each served entity runs on a workload type. Standard CPU gives 4GB per unit of concurrency. If a model needs more memory but no GPU, CPU_MEDIUM and CPU_LARGE give each worker more memory on the same CPU hardware, at the cost of concurrency. The GPU types add accelerators:

    Workload types for custom model serving
    Workload typeHardwareMemory
    CPUCPU4GB per concurrency
    CPU_MEDIUMCPU8GB per concurrency
    CPU_LARGECPU16GB per concurrency
    GPU_SMALL1xT416GB per concurrency
    GPU_MEDIUM1xA10G24GB per concurrency
    MULTIGPU_MEDIUM4xA10G96GB per concurrency
    GPU_MEDIUM_88xA10G192GB per concurrency

    To deploy on a GPU, add a workload_type field to the served entity. Choosing GPU hardware is only half the job: your code must also run predictions on the GPU. MLflow does this automatically for models logged with the PyTorch or Transformers flavors. On GPU endpoints, the number of replicas equals the concurrency value divided by 4. GPU deployments also have extra limits. Container builds take longer than for CPU, and builds plus deployment that run past 60 minutes can time out. GPU autoscaling is also slower than CPU autoscaling. For large language models, the docs point you to Foundation Model APIs instead.

    REST body for a GPU endpoint: workload_type selects the hardwarejson
    {
      "name": "gpu-model-endpoint",
      "config": {
        "served_entities": [{
          "entity_name": "catalog.schema.my-gpu-model",
          "entity_version": "1",
          "workload_type": "GPU_SMALL",
          "workload_size": "Small",
          "scale_to_zero_enabled": false
        }]
      }
    }

    Checkpoint 1 of 5· Check yourself

    A scikit-learn model runs out of memory on the standard CPU workload type. It doesn't benefit from a GPU. What is the most direct fix?

    Sources12

    2.Size concurrency and decide on scale to zero

    Provisioned concurrency is the maximum number of requests an endpoint handles in parallel. Estimate it as QPS × model execution time in seconds. In the API, you set a range with min_provisioned_concurrency and max_provisioned_concurrency. In the UI, you choose a Compute Scale-out size instead:

    Serving UI compute scale-out sizes
    Scale-out sizeConcurrent requests
    Small0-4
    Medium8-16
    Large16-64

    Endpoints scale up almost immediately when traffic rises and scale down every five minutes as it falls. With scale to zero turned on, an endpoint shuts down after 30 minutes of inactivity. The next request then hits a cold start, which usually takes 10-20 seconds but can take minutes, and there is no SLA on it. For GPUs, capacity isn't guaranteed when the endpoint scales back up. That's why the docs say not to use scale to zero for production or latency-sensitive endpoints. Two more limits matter here: requests time out if model computation runs longer than 597 seconds, and route optimization is the recommended option for high-QPS, low-latency workloads.

    Checkpoint 2 of 5· Check yourself

    Which statement about scale to zero on a custom model endpoint is accurate?

    Checkpoint 3 of 5· Exam question

    A served entity hosts a custom regression model behind a production endpoint that receives a steady, moderate volume of prediction requests around the clock, with no long idle periods. Which configuration choice best fits this workload?

    Sources1

    3.Update a live endpoint to a new model version

    After the endpoint is up, you can change most of its configuration, such as the model version or compute, but not its name. You can update it with Edit endpoint in the Serving UI, PUT /api/2.0/serving-endpoints/{name}/config, or the Deployments SDK's update_endpoint_config. The update payload includes served_entities and can include a traffic_config whose routes set each served model's traffic_percentage.

    The config passed to update_endpoint_config: served_entities picks the version, traffic_config routes traffic to itpython
        "served_entities": [
            {
                "entity_name": f"{catalog}.{schema}.{model_name}",
                "entity_version": "1",
                "workload_size": "Small",
                "scale_to_zero_enabled": True
            }
        ],
        "traffic_config": {
            "routes": [
                {
                    "served_model_name": f"{model_name}-1",
                    "traffic_percentage": 100
                }
            ]
        }

    Updates are zero-downtime. The old configuration keeps serving traffic until the new one is ready. Only one update can run at a time, and you can cancel an in-progress update from the Serving UI. If an update fails, the existing configuration stays active, so after each update, check the endpoint status to confirm the change was actually applied. Every update also re-checks the recorded creator's workspace membership and grants. Databricks also runs its own zero-downtime maintenance, during which it reloads models, so your model code must be able to reload at any time.

    Checkpoint 4 of 5· Check yourself

    You submit a configuration that moves an endpoint from model version 3 to version 5. The new container is still building. What happens to incoming prediction requests?

    Checkpoint 5 of 5· Exam question

    An ML engineer wants to compare a new challenger model version against the current production model on the same endpoint, sending most live traffic to the current version while a small share reaches the challenger for evaluation. Which endpoint configuration achieves this?

    Sources12

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Scale to zero is a sensible cost saver for a production endpoint.Why is that wrong?

      While scaled to zero, capacity isn't guaranteed, and the first request has to wait for a cold start. Turn scale to zero off for production and latency-sensitive endpoints.

      Covered in Size concurrency and decide on scale to zero

    2. 2.A failed configuration update leaves the endpoint broken or offline.Why is that wrong?

      If an update fails, the previous active configuration keeps serving as if nothing happened. Check the endpoint status to see whether your change was actually applied.

      Covered in Update a live endpoint to a new model version

    Practise it for real

    Validate a registered custom model, deploy it to a serving endpoint with the MLflow Deployments SDK, and roll it forward to a new version.

    1. 1.Run mlflow.models.predict on the logged model with env_manager="virtualenv" and a sample input.

      Why: This rebuilds the model's logged dependencies in an environment that simulates serving, so dependency problems show up before deployment.

      You should see: Predictions printed to stdout, or a dependency error you fix with update_model_requirements.

    2. 2.Call mlflow.set_registry_uri("databricks-uc"), then get a client with get_deploy_client("databricks").

      Why: The Deployments client accepts the same parameters as the REST API and resolves Unity Catalog model names.

      You should see: A client object, with no output yet.

    3. 3.Call client.create_endpoint with a served_entities entry that uses the full catalog.schema.model entity_name, an entity_version, concurrency values that are multiples of 4, and scale_to_zero_enabled False.

      Why: This creates the endpoint and records your identity as its creator.

      You should see: The endpoint appears in the Serving UI as Not Ready.

    4. 4.Wait for deployment, then check the endpoint state.

      Why: Packaging the model and provisioning the endpoint takes about 10 minutes or more.

      You should see: state.ready is READY, config_update is NOT_UPDATING, and the served entity shows DEPLOYMENT_READY.

    5. 5.Call client.update_endpoint_config with a newer entity_version.

      Why: Updates are zero-downtime, so the old version serves until the new one is ready.

      You should see: config_update shows an update in progress, then the endpoint returns to READY on the new version.

    Stuck? Get a nudge

    If create_endpoint fails with PERMISSION_DENIED, check that your identity has USE CATALOG, USE SCHEMA and EXECUTE on the model.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “When deploying with a GPU, you must make sure that your code is set up so that predictions are run on the GPU”
      ↩︎ Choose a workload type: CPU, larger CPU or GPU
      “provisioned concurrency = queries per second (QPS) * model execution time (s)”
      ↩︎ Size concurrency and decide on scale to zero
      “allows them to scale down to zero after 30 minutes of inactivity”
      ↩︎ Size concurrency and decide on scale to zero
      “During this update process, you are billed for both the old and new endpoint configurations until the transition is complete.”
      ↩︎ Update a live endpoint to a new model version
      “The CPU_MEDIUM and CPU_LARGE workload types let you trade concurrency for more memory per worker on the same CPU hardware.”
      ↩︎ Checkpoint
      “There is no SLA on scale from zero latency.”
      ↩︎ Checkpoint
    2. 2.
      “The number of replicas equals the concurrency value divided by 4.”
      ↩︎ Choose a workload type: CPU, larger CPU or GPU
      “While there is an update in progress, another update cannot be made.”
      ↩︎ Update a live endpoint to a new model version
      “Scale to zero is not recommended for production endpoints, as capacity is not guaranteed when scaled to zero.”
      ↩︎ Exam trap 1
      “Updates to the endpoint configuration can fail. When failures occur the existing active configuration stays effective”
      ↩︎ Exam trap 2
      “Concurrency values must be multiples of 4.”
      ↩︎ Prediction
      “Until the new configuration is ready, the old configuration keeps serving prediction traffic.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.