What you will be able to do
- Choose a CPU or GPU workload type for a custom model
- Size provisioned concurrency from QPS and model execution time
- Decide when scale to zero is appropriate
- Update a live endpoint to a new model version without downtime
1.Choose a workload type: CPU, larger CPU or GPU
Each served entity runs on a workload type. Standard CPU gives 4GB per unit of concurrency. If a model needs more memory but no GPU, CPU_MEDIUM and CPU_LARGE give each worker more memory on the same CPU hardware, at the cost of concurrency. The GPU types add accelerators:
| Workload type | Hardware | Memory |
|---|---|---|
| CPU | CPU | 4GB per concurrency |
| CPU_MEDIUM | CPU | 8GB per concurrency |
| CPU_LARGE | CPU | 16GB per concurrency |
| GPU_SMALL | 1xT4 | 16GB per concurrency |
| GPU_MEDIUM | 1xA10G | 24GB per concurrency |
| MULTIGPU_MEDIUM | 4xA10G | 96GB per concurrency |
| GPU_MEDIUM_8 | 8xA10G | 192GB per concurrency |
To deploy on a GPU, add a workload_type field to the served entity. Choosing GPU hardware is only half the job: your code must also run predictions on the GPU. MLflow does this automatically for models logged with the PyTorch or Transformers flavors. On GPU endpoints, the number of replicas equals the concurrency value divided by 4. GPU deployments also have extra limits. Container builds take longer than for CPU, and builds plus deployment that run past 60 minutes can time out. GPU autoscaling is also slower than CPU autoscaling. For large language models, the docs point you to Foundation Model APIs instead.
{
"name": "gpu-model-endpoint",
"config": {
"served_entities": [{
"entity_name": "catalog.schema.my-gpu-model",
"entity_version": "1",
"workload_type": "GPU_SMALL",
"workload_size": "Small",
"scale_to_zero_enabled": false
}]
}
}Checkpoint 1 of 5· Check yourself
A scikit-learn model runs out of memory on the standard CPU workload type. It doesn't benefit from a GPU. What is the most direct fix?
CPU_MEDIUM and CPU_LARGE give each worker more memory on the same CPU hardware. Raising concurrency adds parallel capacity, not memory per request.
“The CPU_MEDIUM and CPU_LARGE workload types let you trade concurrency for more memory per worker on the same CPU hardware.”Source: docs.databricks.com
2.Size concurrency and decide on scale to zero
Provisioned concurrency is the maximum number of requests an endpoint handles in parallel. Estimate it as QPS × model execution time in seconds. In the API, you set a range with min_provisioned_concurrency and max_provisioned_concurrency. In the UI, you choose a Compute Scale-out size instead:
| Scale-out size | Concurrent requests |
|---|---|
| Small | 0-4 |
| Medium | 8-16 |
| Large | 16-64 |
Endpoints scale up almost immediately when traffic rises and scale down every five minutes as it falls. With scale to zero turned on, an endpoint shuts down after 30 minutes of inactivity. The next request then hits a cold start, which usually takes 10-20 seconds but can take minutes, and there is no SLA on it. For GPUs, capacity isn't guaranteed when the endpoint scales back up. That's why the docs say not to use scale to zero for production or latency-sensitive endpoints. Two more limits matter here: requests time out if model computation runs longer than 597 seconds, and route optimization is the recommended option for high-QPS, low-latency workloads.
Checkpoint 2 of 5· Check yourself
Which statement about scale to zero on a custom model endpoint is accurate?
Scale to zero kicks in after 30 minutes idle. The cold start usually takes 10-20 seconds, can take minutes, and has no SLA.
“There is no SLA on scale from zero latency.”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
A served entity hosts a custom regression model behind a production endpoint that receives a steady, moderate volume of prediction requests around the clock, with no long idle periods. Which configuration choice best fits this workload?
Correct answer: A — Set `workload_size` to Medium and leave `scale_to_zero_enabled` set to false, since capacity is not guaranteed once an endpoint scales down and this traffic pattern never truly goes idle.
- A. Sizing the served entity at Medium provisions enough capacity for steady moderate traffic, and disabling scale-to-zero avoids the unguaranteed cold-start capacity risk that scale-to-zero introduces for an endpoint that is effectively always busy.
- B. Enabling scale-to-zero is meant for endpoints with genuine idle windows; against continuous traffic it repeatedly tears down and reprovisions compute, adding cold-start latency instead of saving cost.
- C. GPU workload types accelerate deep-learning-style matrix operations, but a scikit-learn regression model does not use the GPU compute path, so this setting adds cost without improving latency.
- D. Pinning minimum and maximum provisioned concurrency to the same fixed value removes autoscaling headroom, so the endpoint cannot absorb bursts above that fixed level without throttling requests.
Sources1
3.Update a live endpoint to a new model version
After the endpoint is up, you can change most of its configuration, such as the model version or compute, but not its name. You can update it with Edit endpoint in the Serving UI, PUT /api/2.0/serving-endpoints/{name}/config, or the Deployments SDK's update_endpoint_config. The update payload includes served_entities and can include a traffic_config whose routes set each served model's traffic_percentage.
"served_entities": [
{
"entity_name": f"{catalog}.{schema}.{model_name}",
"entity_version": "1",
"workload_size": "Small",
"scale_to_zero_enabled": True
}
],
"traffic_config": {
"routes": [
{
"served_model_name": f"{model_name}-1",
"traffic_percentage": 100
}
]
}Updates are zero-downtime. The old configuration keeps serving traffic until the new one is ready. Only one update can run at a time, and you can cancel an in-progress update from the Serving UI. If an update fails, the existing configuration stays active, so after each update, check the endpoint status to confirm the change was actually applied. Every update also re-checks the recorded creator's workspace membership and grants. Databricks also runs its own zero-downtime maintenance, during which it reloads models, so your model code must be able to reload at any time.
Both of them. You pay for the old and the new configuration until the switch-over finishes.
Checkpoint 4 of 5· Check yourself
You submit a configuration that moves an endpoint from model version 3 to version 5. The new container is still building. What happens to incoming prediction requests?
Model Serving keeps the existing configuration live until the new one is ready, so updates cause no downtime.
“Until the new configuration is ready, the old configuration keeps serving prediction traffic.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
An ML engineer wants to compare a new challenger model version against the current production model on the same endpoint, sending most live traffic to the current version while a small share reaches the challenger for evaluation. Which endpoint configuration achieves this?
Correct answer: A — Define both model versions as served entities on one endpoint, then add a `traffic_config` with routes that send most traffic to the current entity and the rest to the challenger entity.
- A. Listing both versions as served entities and adding a traffic_config with per-route traffic percentages is exactly how Model Serving performs canary and A/B routing, letting most requests hit the proven version while a small share reaches the challenger.
- B. Two independent endpoints do not share traffic-splitting logic, so the challenger never receives any of the live traffic being compared, which defeats the purpose of an A/B test on the same endpoint.
- C. Serving only the challenger sends it all production traffic immediately with no comparison against the current version, which is the opposite of a controlled traffic split.
- D. Without a traffic_config block, Model Serving does not default to an even split between served entities, so traffic distribution between the two versions would be undefined rather than intentional.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Scale to zero is a sensible cost saver for a production endpoint.Why is that wrong?
While scaled to zero, capacity isn't guaranteed, and the first request has to wait for a cold start. Turn scale to zero off for production and latency-sensitive endpoints.
Covered in Size concurrency and decide on scale to zero
2.A failed configuration update leaves the endpoint broken or offline.Why is that wrong?
If an update fails, the previous active configuration keeps serving as if nothing happened. Check the endpoint status to see whether your change was actually applied.
Practise it for real
Validate a registered custom model, deploy it to a serving endpoint with the MLflow Deployments SDK, and roll it forward to a new version.
1.Run mlflow.models.predict on the logged model with env_manager="virtualenv" and a sample input.
Why: This rebuilds the model's logged dependencies in an environment that simulates serving, so dependency problems show up before deployment.
You should see: Predictions printed to stdout, or a dependency error you fix with update_model_requirements.
2.Call mlflow.set_registry_uri("databricks-uc"), then get a client with get_deploy_client("databricks").
Why: The Deployments client accepts the same parameters as the REST API and resolves Unity Catalog model names.
You should see: A client object, with no output yet.
3.Call client.create_endpoint with a served_entities entry that uses the full catalog.schema.model entity_name, an entity_version, concurrency values that are multiples of 4, and scale_to_zero_enabled False.
Why: This creates the endpoint and records your identity as its creator.
You should see: The endpoint appears in the Serving UI as Not Ready.
4.Wait for deployment, then check the endpoint state.
Why: Packaging the model and provisioning the endpoint takes about 10 minutes or more.
You should see: state.ready is READY, config_update is NOT_UPDATING, and the served entity shows DEPLOYMENT_READY.
5.Call client.update_endpoint_config with a newer entity_version.
Why: Updates are zero-downtime, so the old version serves until the new one is ready.
You should see: config_update shows an update in progress, then the endpoint returns to READY on the new version.
Stuck? Get a nudge
If create_endpoint fails with PERMISSION_DENIED, check that your identity has USE CATALOG, USE SCHEMA and EXECUTE on the model.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“When deploying with a GPU, you must make sure that your code is set up so that predictions are run on the GPU”
↩︎ Choose a workload type: CPU, larger CPU or GPU“provisioned concurrency = queries per second (QPS) * model execution time (s)”
↩︎ Size concurrency and decide on scale to zero“allows them to scale down to zero after 30 minutes of inactivity”
↩︎ Size concurrency and decide on scale to zero“During this update process, you are billed for both the old and new endpoint configurations until the transition is complete.”
↩︎ Update a live endpoint to a new model version“The CPU_MEDIUM and CPU_LARGE workload types let you trade concurrency for more memory per worker on the same CPU hardware.”
↩︎ Checkpoint“There is no SLA on scale from zero latency.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/machine-learning/model-serving/create-manage-serving-endpointsOfficial docs
“The number of replicas equals the concurrency value divided by 4.”
↩︎ Choose a workload type: CPU, larger CPU or GPU“While there is an update in progress, another update cannot be made.”
↩︎ Update a live endpoint to a new model version“Scale to zero is not recommended for production endpoints, as capacity is not guaranteed when scaled to zero.”
↩︎ Exam trap 1“Updates to the endpoint configuration can fail. When failures occur the existing active configuration stays effective”
↩︎ Exam trap 2“Concurrency values must be multiples of 4.”
↩︎ Prediction“Until the new configuration is ready, the old configuration keeps serving prediction traffic.”
↩︎ Checkpoint