CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 4 · Lesson 47/48

    Deploy a Custom Model to a Databricks Model Serving Endpoint

    Deploy and query a model for realtime inference

    11 min read
    2.08% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Describe the path from a logged MLflow model to a live real-time serving endpoint
    • Identify the creator identity and Unity Catalog grants an endpoint needs
    • Create a custom model serving endpoint with the Serving UI, the REST API, or the MLflow Deployments SDK
    • Choose scale-out size, scale-to-zero, and workload type for a served entity

    Key concept

    Model serving endpoint — A serverless, autoscaling deployment that puts one or more registered MLflow models behind a REST API, so applications can send scoring requests and get predictions back in real time.

    1.From a logged model to a REST endpoint

    Databricks Model Serving is the platform's service for deploying models for real-time serving and batch inference. It runs on serverless compute, scales up or down as demand changes, and exposes every served model as a REST API. You manage it through one UI, a REST API, and the MLflow Deployments API, which all cover the same create, read, update and delete (CRUD) and query tasks.

    This lesson covers custom models. Databricks uses that term for traditional ML models or customized Python models packaged in MLflow format, such as scikit-learn, XGBoost, PyTorch and Hugging Face transformer models. Foundation models and external models are served differently and are not covered here.

    Deployment always follows the same path. First you log the model in MLflow format. You can rely on autologging, call a built-in flavor such as mlflow.sklearn.log_model, or write a custom pyfunc model when you need arbitrary Python code or extra steps before or after inference. Next you register the model, ideally in Unity Catalog, though the legacy workspace model registry also works. Only then can you create an endpoint and query it.

    Logging a signature and an input example is recommended for every model. For Unity Catalog, the signature is required. The input example is useful later too: the Serving UI can load it as a sample request.

    Checkpoint 1 of 6· Put it in order

    Put the steps for serving a custom model in order.

    1. 1.Send scoring requests to the endpoint
    2. 2.Create a model serving endpoint that serves the registered version
    3. 3.Log the model in MLflow format (built-in flavor or pyfunc)
    4. 4.Register the logged model in Unity Catalog

    Sources12

    2.Who owns the endpoint: creator identity and grants

    When you create an endpoint, Databricks records the identity that made the call as the endpoint's creator, and that identity is usually a service principal. The endpoint uses it to reach Unity Catalog resources, and you cannot change it after creation.

    For each served Unity Catalog model, the recorded creator needs USE CATALOG on the catalog, USE SCHEMA on the schema, and EXECUTE on the model. Databricks checks these grants when the endpoint is created and again whenever it is updated. If a grant is missing, the request fails with PERMISSION_DENIED. Updates also fail if the recorded creator has left the workspace, even when the person making the update has valid permissions. In that case, the only fix is to delete the endpoint and recreate it under a service principal that is still a workspace member. For that reason, Databricks advises using a long-lived service principal owned by your team as the creator, not a personal user account.

    Checkpoint 2 of 6· Check yourself

    An engineer created an endpoint with their personal account and then left the company. A teammate with CAN_MANAGE now tries to update the served model version. What happens?

    Sources3

    3.Creating an endpoint in the Serving UI

    In the UI, click Serving in the sidebar, then Create serving endpoint. Give the endpoint a name. Names cannot start with databricks-, because that prefix is reserved for preconfigured endpoints. Then configure a served entity:

    - Choose *My models- Unity Catalog* or *My models- Model Registry*, then pick the model and the version. - Set the percentage of traffic that goes to this served model. - Pick a compute type (CPU or GPU). - Pick a Compute Scale-out size. This is how many requests the model can process at the same time: Small handles 0–4, Medium 8–16, and Large 16–64. Size it at roughly QPS × model run time. - Decide whether the endpoint should scale to zero when idle.

    You can add more served entities to the same endpoint and split traffic between them. You can also turn on route optimization, which is recommended for endpoints with high QPS and throughput requirements.

    When you click Create, the Serving endpoints page shows the endpoint as *Not Ready* while Databricks builds it. Wait for it to become ready before you send traffic.

    Checkpoint 3 of 6· Check yourself

    A model takes about 0.5 seconds per request and must handle about 16 queries per second. Which Compute Scale-out size fits best?

    Checkpoint 4 of 6· Exam question

    A fraud-detection model must return a prediction in well under 100ms for each individual transaction as a mobile app submits it, and the app cannot wait for a scheduled job to run. Which deployment approach satisfies this requirement?

    Sources3

    4.Creating an endpoint with the REST API or MLflow Deployments SDK

    To create an endpoint in code, send POST /api/2.0/serving-endpoints. The request body holds a config with a list of served_entities. Each entity names the Unity Catalog model by its full three-level name (catalog.schema.model) and gives a version. You can let Databricks size compute with workload_size, or set your own concurrency with min_provisioned_concurrency and max_provisioned_concurrency. Concurrency values must be multiples of 4.

    REST body that creates an endpoint serving version 3 of a Unity Catalog model, with custom concurrencyjson
    {
      "name": "uc-model-endpoint",
      "config":
      {
        "served_entities": [
          {
            "name": "ads-entity",
            "entity_name": "catalog.schema.my-ads-model",
            "entity_version": "3",
            "min_provisioned_concurrency": 4,
            "max_provisioned_concurrency": 12,
            "scale_to_zero_enabled": false
          }
        ]
      }
    }

    The endpoint is ready when the response shows state.ready as READY and config_update as NOT_UPDATING, and when the served entity's deployment state is DEPLOYMENT_READY.

    The MLflow Deployments SDK accepts the same parameters as the REST API. You get a client with get_deploy_client("databricks"), point MLflow at the Unity Catalog registry, and pass the same config dictionary. The Databricks Workspace Client SDK is a third option. It has typed classes, EndpointCoreConfigInput and ServedEntityInput, for the same fields.

    Checkpoint 5 of 6· Fill the gap

    Which MLflow Deployments client method creates a new serving endpoint?

    import mlflow
    from mlflow.deployments import get_deploy_client
    
    mlflow.set_registry_uri("databricks-uc")
    client = get_deploy_client("databricks")
    
    endpoint = client. ? (
        name="unity-catalog-model-endpoint",

    Sources3

    5.Compute types and the serving container

    Every served entity runs on a workload type. Standard CPU gives 4GB of memory per unit of concurrency. CPU_MEDIUM and CPU_LARGE give each worker more memory on the same CPU hardware, at the cost of concurrency. GPU types use the workload_type field. On GPU endpoints, the number of replicas equals the concurrency value divided by 4.

    Model Serving workload types and memory
    Workload typeGPU instanceMemory
    CPU—4GB per concurrency
    CPU_MEDIUM—8GB per concurrency
    CPU_LARGE—16GB per concurrency
    GPU_SMALL1xT416GB per concurrency
    GPU_MEDIUM1xA10G24GB per concurrency
    MULTIGPU_MEDIUM4xA10G96GB per concurrency

    Checkpoint 6 of 6· Check yourself

    A GPU endpoint is created with min_provisioned_concurrency set to 16. How many replicas does it provision?

    During deployment, Databricks builds a production-grade container from the MLflow model. Built-in flavors capture their package dependencies automatically. For custom pyfunc models, you list dependencies yourself with pip_requirements, conda_env or extra_pip_requirements. If a dependency is missing, the deployment fails with a dependency error, so Databricks recommends testing the model locally first. Model Serving also does not patch existing model images: the latest security patches only arrive in an image built from a new model version.

    Sources21

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Turning on scale to zero for production endpoints is a free cost saving.Why is that wrong?

      Databricks does not recommend scale to zero for production. Capacity is not guaranteed, and scaling back up adds cold-start latency.

      Covered in Creating an endpoint in the Serving UI

    2. 2.If the endpoint creator leaves, an admin can reassign the endpoint to someone else.Why is that wrong?

      The recorded creator identity is fixed when the endpoint is created. If it loses its grants or workspace membership, you must delete the endpoint and recreate it under a valid service principal.

      Covered in Who owns the endpoint: creator identity and grants

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Model Serving offers a unified REST API and MLflow Deployment API for CRUD and querying tasks.”
      ↩︎ From a logged model to a REST endpoint
      “A new model image created from a new model version will contain the latest patches.”
      ↩︎ Compute types and the serving container
      “Each model you serve is available as a REST API that you can integrate into your web or client application.”
      ↩︎ Key concept
    2. 2.
      “Signatures are necessary for logging models to the Unity Catalog.”
      ↩︎ From a logged model to a REST endpoint
      “During deployment, a production-grade container is built and deployed as the endpoint.”
      ↩︎ Compute types and the serving container
      “After the model is logged, register it in the Unity Catalog (recommended) or the workspace registry.”
      ↩︎ Checkpoint
    3. 3.
      “USE CATALOG on the catalog, USE SCHEMA on the schema, EXECUTE on the model”
      ↩︎ Who owns the endpoint: creator identity and grants
      “Endpoint names cannot use the databricks- prefix.”
      ↩︎ Creating an endpoint in the Serving UI
      “Concurrency values must be multiples of 4.”
      ↩︎ Creating an endpoint with the REST API or MLflow Deployments SDK
      “The endpoint's config_update state is NOT_UPDATING and the served model is in a READY state.”
      ↩︎ Creating an endpoint with the REST API or MLflow Deployments SDK
      “Scale to zero is not recommended for production endpoints, as capacity is not guaranteed when scaled to zero.”
      ↩︎ Exam trap 1
      “is used to access Unity Catalog resources on behalf of the endpoint and cannot be changed after creation.”
      ↩︎ Exam trap 2
      “Updates fail with PERMISSION_DENIED if the recorded creator is no longer a workspace member, even when the caller has valid permissions.”
      ↩︎ Checkpoint
      “Scale to zero is not recommended for production endpoints, as capacity is not guaranteed when scaled to zero.”
      ↩︎ Prediction
      “This number should be roughly equal to QPS x model run time.”
      ↩︎ Checkpoint
      “The number of replicas equals the concurrency value divided by 4.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Query a Real-Time Model Serving Endpoint

    Spotted a mistake, or was something unclear? Tell us.