What you will be able to do
- Describe the path from a logged MLflow model to a live real-time serving endpoint
- Identify the creator identity and Unity Catalog grants an endpoint needs
- Create a custom model serving endpoint with the Serving UI, the REST API, or the MLflow Deployments SDK
- Choose scale-out size, scale-to-zero, and workload type for a served entity
Key concept
Model serving endpoint — A serverless, autoscaling deployment that puts one or more registered MLflow models behind a REST API, so applications can send scoring requests and get predictions back in real time.
1.From a logged model to a REST endpoint
Databricks Model Serving is the platform's service for deploying models for real-time serving and batch inference. It runs on serverless compute, scales up or down as demand changes, and exposes every served model as a REST API. You manage it through one UI, a REST API, and the MLflow Deployments API, which all cover the same create, read, update and delete (CRUD) and query tasks.
This lesson covers custom models. Databricks uses that term for traditional ML models or customized Python models packaged in MLflow format, such as scikit-learn, XGBoost, PyTorch and Hugging Face transformer models. Foundation models and external models are served differently and are not covered here.
Deployment always follows the same path. First you log the model in MLflow format. You can rely on autologging, call a built-in flavor such as mlflow.sklearn.log_model, or write a custom pyfunc model when you need arbitrary Python code or extra steps before or after inference. Next you register the model, ideally in Unity Catalog, though the legacy workspace model registry also works. Only then can you create an endpoint and query it.
Logging a signature and an input example is recommended for every model. For Unity Catalog, the signature is required. The input example is useful later too: the Serving UI can load it as a sample request.
Checkpoint 1 of 6· Put it in order
Put the steps for serving a custom model in order.
- 1.Send scoring requests to the endpoint
- 2.Create a model serving endpoint that serves the registered version
- 3.Log the model in MLflow format (built-in flavor or pyfunc)
- 4.Register the logged model in Unity Catalog
An endpoint serves a registered model version, so you log and register first, and you can only query once the endpoint exists.
“After the model is logged, register it in the Unity Catalog (recommended) or the workspace registry.”Source: docs.databricks.com
2.Who owns the endpoint: creator identity and grants
When you create an endpoint, Databricks records the identity that made the call as the endpoint's creator, and that identity is usually a service principal. The endpoint uses it to reach Unity Catalog resources, and you cannot change it after creation.
For each served Unity Catalog model, the recorded creator needs USE CATALOG on the catalog, USE SCHEMA on the schema, and EXECUTE on the model. Databricks checks these grants when the endpoint is created and again whenever it is updated. If a grant is missing, the request fails with PERMISSION_DENIED. Updates also fail if the recorded creator has left the workspace, even when the person making the update has valid permissions. In that case, the only fix is to delete the endpoint and recreate it under a service principal that is still a workspace member. For that reason, Databricks advises using a long-lived service principal owned by your team as the creator, not a personal user account.
Checkpoint 2 of 6· Check yourself
An engineer created an endpoint with their personal account and then left the company. A teammate with CAN_MANAGE now tries to update the served model version. What happens?
Updates check the recorded creator's membership, not just the caller's. The creator cannot be changed, so the endpoint has to be recreated under a current service principal.
“Updates fail with PERMISSION_DENIED if the recorded creator is no longer a workspace member, even when the caller has valid permissions.”Source: docs.databricks.com
Sources3
3.Creating an endpoint in the Serving UI
In the UI, click Serving in the sidebar, then Create serving endpoint. Give the endpoint a name. Names cannot start with databricks-, because that prefix is reserved for preconfigured endpoints. Then configure a served entity:
- Choose *My models- Unity Catalog* or *My models- Model Registry*, then pick the model and the version. - Set the percentage of traffic that goes to this served model. - Pick a compute type (CPU or GPU). - Pick a Compute Scale-out size. This is how many requests the model can process at the same time: Small handles 0–4, Medium 8–16, and Large 16–64. Size it at roughly QPS × model run time. - Decide whether the endpoint should scale to zero when idle.
You can add more served entities to the same endpoint and split traffic between them. You can also turn on route optimization, which is recommended for endpoints with high QPS and throughput requirements.
When you click Create, the Serving endpoints page shows the endpoint as *Not Ready* while Databricks builds it. Wait for it to become ready before you send traffic.
Checkpoint 3 of 6· Check yourself
A model takes about 0.5 seconds per request and must handle about 16 queries per second. Which Compute Scale-out size fits best?
16 QPS × 0.5 s gives about 8 concurrent requests, which falls in the Medium range.
“This number should be roughly equal to QPS x model run time.”Source: docs.databricks.com
Checkpoint 4 of 6· Exam question
A fraud-detection model must return a prediction in well under 100ms for each individual transaction as a mobile app submits it, and the app cannot wait for a scheduled job to run. Which deployment approach satisfies this requirement?
Correct answer: A — Deploy the registered model behind a Databricks Model Serving endpoint and have the app send each transaction as a synchronous REST API request to that endpoint.
- A. A Model Serving endpoint hosts the model behind a REST API that responds to a single request in real time, which is what a sub-100ms per-transaction latency requirement from an interactive app needs.
- B. A Lakeflow Spark Declarative Pipeline processes transactions on a micro-batch trigger and writes results to a table, which adds pipeline scheduling latency that an interactive mobile request cannot tolerate.
- C. Loading the model in a scheduled notebook and scoring accumulated rows is batch inference, so results only become available on the job's schedule rather than synchronously for each transaction.
- D. A vectorized pandas UDF run during an ETL job scores data already collected in a DataFrame, which again ties predictions to a job schedule instead of returning a result for one transaction immediately.
Sources3
4.Creating an endpoint with the REST API or MLflow Deployments SDK
To create an endpoint in code, send POST /api/2.0/serving-endpoints. The request body holds a config with a list of served_entities. Each entity names the Unity Catalog model by its full three-level name (catalog.schema.model) and gives a version. You can let Databricks size compute with workload_size, or set your own concurrency with min_provisioned_concurrency and max_provisioned_concurrency. Concurrency values must be multiples of 4.
{
"name": "uc-model-endpoint",
"config":
{
"served_entities": [
{
"name": "ads-entity",
"entity_name": "catalog.schema.my-ads-model",
"entity_version": "3",
"min_provisioned_concurrency": 4,
"max_provisioned_concurrency": 12,
"scale_to_zero_enabled": false
}
]
}
}The endpoint is ready when the response shows state.ready as READY and config_update as NOT_UPDATING, and when the served entity's deployment state is DEPLOYMENT_READY.
The MLflow Deployments SDK accepts the same parameters as the REST API. You get a client with get_deploy_client("databricks"), point MLflow at the Unity Catalog registry, and pass the same config dictionary. The Databricks Workspace Client SDK is a third option. It has typed classes, EndpointCoreConfigInput and ServedEntityInput, for the same fields.
Checkpoint 5 of 6· Fill the gap
Which MLflow Deployments client method creates a new serving endpoint?
import mlflow
from mlflow.deployments import get_deploy_client
mlflow.set_registry_uri("databricks-uc")
client = get_deploy_client("databricks")
endpoint = client. ? (
name="unity-catalog-model-endpoint",create_endpoint creates the endpoint. update_endpoint_config changes an existing endpoint, and predict sends scoring requests.
Sources3
5.Compute types and the serving container
Every served entity runs on a workload type. Standard CPU gives 4GB of memory per unit of concurrency. CPU_MEDIUM and CPU_LARGE give each worker more memory on the same CPU hardware, at the cost of concurrency. GPU types use the workload_type field. On GPU endpoints, the number of replicas equals the concurrency value divided by 4.
| Workload type | GPU instance | Memory |
|---|---|---|
| CPU | — | 4GB per concurrency |
| CPU_MEDIUM | — | 8GB per concurrency |
| CPU_LARGE | — | 16GB per concurrency |
| GPU_SMALL | 1xT4 | 16GB per concurrency |
| GPU_MEDIUM | 1xA10G | 24GB per concurrency |
| MULTIGPU_MEDIUM | 4xA10G | 96GB per concurrency |
Checkpoint 6 of 6· Check yourself
A GPU endpoint is created with min_provisioned_concurrency set to 16. How many replicas does it provision?
On GPU endpoints replicas equal concurrency divided by 4, so 16 gives 4 replicas.
“The number of replicas equals the concurrency value divided by 4.”Source: docs.databricks.com
During deployment, Databricks builds a production-grade container from the MLflow model. Built-in flavors capture their package dependencies automatically. For custom pyfunc models, you list dependencies yourself with pip_requirements, conda_env or extra_pip_requirements. If a dependency is missing, the deployment fails with a dependency error, so Databricks recommends testing the model locally first. Model Serving also does not patch existing model images: the latest security patches only arrive in an image built from a new model version.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Turning on scale to zero for production endpoints is a free cost saving.Why is that wrong?
Databricks does not recommend scale to zero for production. Capacity is not guaranteed, and scaling back up adds cold-start latency.
Covered in Creating an endpoint in the Serving UI
2.If the endpoint creator leaves, an admin can reassign the endpoint to someone else.Why is that wrong?
The recorded creator identity is fixed when the endpoint is created. If it loses its grants or workspace membership, you must delete the endpoint and recreate it under a valid service principal.
Covered in Who owns the endpoint: creator identity and grants
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Model Serving offers a unified REST API and MLflow Deployment API for CRUD and querying tasks.”
↩︎ From a logged model to a REST endpoint“A new model image created from a new model version will contain the latest patches.”
↩︎ Compute types and the serving container“Each model you serve is available as a REST API that you can integrate into your web or client application.”
↩︎ Key concept - 2.
“Signatures are necessary for logging models to the Unity Catalog.”
↩︎ From a logged model to a REST endpoint“During deployment, a production-grade container is built and deployed as the endpoint.”
↩︎ Compute types and the serving container“After the model is logged, register it in the Unity Catalog (recommended) or the workspace registry.”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/machine-learning/model-serving/create-manage-serving-endpointsOfficial docs
“USE CATALOG on the catalog, USE SCHEMA on the schema, EXECUTE on the model”
↩︎ Who owns the endpoint: creator identity and grants“Endpoint names cannot use the databricks- prefix.”
↩︎ Creating an endpoint in the Serving UI“Concurrency values must be multiples of 4.”
↩︎ Creating an endpoint with the REST API or MLflow Deployments SDK“The endpoint's config_update state is NOT_UPDATING and the served model is in a READY state.”
↩︎ Creating an endpoint with the REST API or MLflow Deployments SDK“Scale to zero is not recommended for production endpoints, as capacity is not guaranteed when scaled to zero.”
↩︎ Exam trap 1“is used to access Unity Catalog resources on behalf of the endpoint and cannot be changed after creation.”
↩︎ Exam trap 2“Updates fail with PERMISSION_DENIED if the recorded creator is no longer a workspace member, even when the caller has valid permissions.”
↩︎ Checkpoint“Scale to zero is not recommended for production endpoints, as capacity is not guaranteed when scaled to zero.”
↩︎ Prediction“This number should be roughly equal to QPS x model run time.”
↩︎ Checkpoint“The number of replicas equals the concurrency value divided by 4.”
↩︎ Checkpoint