CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 4 · Lesson 44/48

    Deploy a Custom MLflow Model to a Databricks Model Serving Endpoint

    Deploy a custom model to a model endpoint

    11 min read
    2.08% of exam
    4 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Describe the log, register, then create-endpoint path for a custom model
    • Declare and validate model dependencies before deployment
    • Create a custom model serving endpoint with the Serving UI, the REST API or the MLflow Deployments SDK
    • Explain which identity and Unity Catalog grants an endpoint depends on

    Key concept

    Custom model serving endpoint — A custom model is any Python model or code packaged in MLflow format and registered in Unity Catalog or the workspace registry. A serving endpoint runs a specific registered version of that model in a container that Databricks builds, and exposes it as a REST API.

    1.From trained model to servable custom model

    Databricks Model Serving turns a registered model into a REST API. It runs on serverless compute that scales up or down as demand changes. A custom model is any Python model or custom code packaged in the MLflow format. Examples include scikit-learn, XGBoost, PyTorch and Hugging Face transformers models. Foundation models served through Foundation Model APIs and external models are separate categories.

    The deployment path has three steps. First, log the model in MLflow format, using either a built-in MLflow flavor or pyfunc. Second, register it in Unity Catalog (recommended) or in the legacy workspace model registry. Third, create a serving endpoint that points at a specific version of that registered model. To register a model in Unity Catalog, you must log it with a signature. Databricks also recommends logging an input example, because the Serving UI can later load it into a test request with Show Example.

    The three ways to log a model for Model Serving
    Logging techniqueWhen to use it
    AutologgingTurned on automatically in Databricks Runtime for ML. It is the easiest option but gives you less control.
    Built-in flavors (e.g. mlflow.sklearn.log_model)Log the model manually when you want more detailed control
    Custom logging with pyfunc (mlflow.pyfunc.PythonModel)Use for arbitrary Python code, or when you need extra steps before or after inference

    Checkpoint 1 of 6· Put it in order

    Put the steps for deploying a custom model in order

    1. 1.Register the logged model in Unity Catalog (recommended) or the workspace registry
    2. 2.Create a model serving endpoint for the registered model version
    3. 3.Log the model in MLflow format with a built-in flavor or pyfunc

    Sources1

    2.Package dependencies and validate before you deploy

    During deployment, Databricks builds a production-grade container for the endpoint. The container only includes libraries that MLflow captured automatically or that you specified when you logged the model. The base image may provide some system-level packages, but you must declare application-level dependencies in the model yourself. If a dependency is missing, you get errors during deployment.

    For native flavors, MLflow captures package dependencies automatically. For custom pyfunc models, you add them explicitly with these logging parameters:

    Logging parameters that control what goes into the serving container
    ParameterWhat it does
    pip_requirementsSets the full list of pip packages for the model
    conda_envDefines the environment as a conda specification (channels, dependencies)
    extra_pip_requirementsAdds packages on top of the ones captured automatically
    code_pathPackages local code files, such as helper modules, with the model

    Building an endpoint can take about 10 minutes, so Databricks recommends validating the model first instead of waiting for a deployment to fail. mlflow.models.predict takes the model_uri and some input data. With the default env_manager="virtualenv", it rebuilds the model's logged dependencies in an isolated environment that simulates the serving runtime. A local option exists, but the docs call it potentially error prone for serving validation. pip_requirements_override lets you try different package versions without logging the model again.

    Testing a logged model in a virtualenv that simulates servingpython
    mlflow.models.predict(
      model_uri=model_uri,
      input_data={"col1": 34.2, "col2": 11.2, "col3": "green"},
      content_type="json",
      env_manager="virtualenv",
      install_mlflow=False,
      pip_requirements_override=["pillow==10.3.0", "scipy==1.13.0"],
    )

    If a requirement turns out to be wrong, update_model_requirements() edits the pip_requirements.txt file inside the model artifact in place, so you don't have to log a new model. Serving endpoints also expect a specific JSON input format. validate_serving_input checks a serving payload against the logged model before you deploy.

    Checkpoint 2 of 6· Check yourself

    You logged a model with mlflow.sklearn.log_model. MLflow captured scikit-learn automatically, but the model also needs one extra package. Which parameter adds it without replacing the captured list?

    Sources12

    3.Create the endpoint: Serving UI, REST API or MLflow Deployments SDK

    You can create a custom model endpoint in three ways: the Serving UI, the REST API or the MLflow Deployments SDK. The Databricks Workspace Client SDK is a fourth option.

    Serving UI. Click Serving in the sidebar, then Create serving endpoint. Enter a name. Names can't start with databricks-, because that prefix is reserved for preconfigured endpoints. Under Served entities, choose *My models- Unity Catalog* or *My models- Model Registry*. Then pick the model and version, the percentage of traffic it receives, the compute type, the compute scale-out, and whether the endpoint can scale to zero. Click Create. The endpoint's state shows as Not Ready while it builds.

    REST API. Send POST /api/2.0/serving-endpoints. Each item in served_entities names a model and a version. For a Unity Catalog model, entity_name must be the full catalog.schema.model name.

    REST request body that serves version 3 of a Unity Catalog modeljson
    {
      "name": "uc-model-endpoint",
      "config":
      {
        "served_entities": [
          {
            "name": "ads-entity",
            "entity_name": "catalog.schema.my-ads-model",
            "entity_version": "3",
            "min_provisioned_concurrency": 4,
            "max_provisioned_concurrency": 12,
            "scale_to_zero_enabled": false
          }
        ]
      }
    }

    MLflow Deployments SDK. mlflow.deployments.get_deploy_client("databricks") returns a client whose create, update and delete methods accept the same parameters as the REST API. Set the registry URI to databricks-uc first so the model name resolves in Unity Catalog.

    Checkpoint 3 of 6· Fill the gap

    Which client method completes this MLflow Deployments SDK sample?

    mlflow.set_registry_uri("databricks-uc")
    client = get_deploy_client("databricks")
    
    endpoint = client. ? (
        name="unity-catalog-model-endpoint",
        config={
            "served_entities": [
                {
                    "name": "ads-entity",
                    "entity_name": "catalog.schema.my-ads-model",
                    "entity_version": "3",
                    "min_provisioned_concurrency": 4,
                    "max_provisioned_concurrency": 12,
                    "scale_to_zero_enabled": False
                }
            ]
        }
    )

    The endpoint is live once the response shows state.ready as READY and config_update as NOT_UPDATING, with the served entity's deployment state at DEPLOYMENT_READY. To test it, open the endpoint page in the Serving UI, choose Query endpoint, paste JSON input, and click Send Request.

    Checkpoint 4 of 6· Exam question

    A team has trained a custom scikit-learn fraud-detection model and registered it in Unity Catalog. The application must return a prediction within 50 milliseconds of each transaction as it happens, and transactions arrive continuously throughout the day. Which approach should the team use to serve predictions?

    Checkpoint 5 of 6· Check yourself

    A team tries to create an endpoint named databricks-churn-scorer for their Unity Catalog model. What happens?

    Sources34

    4.The endpoint's recorded creator and its Unity Catalog grants

    When you create an endpoint, Databricks records the calling identity as the endpoint's creator. The endpoint uses that identity, usually a service principal, to access Unity Catalog resources, and you can't change it later. Both the caller and the recorded creator must be workspace members and hold the workspace-access entitlement.

    For each served Unity Catalog model, the creator needs USE CATALOG on the catalog, USE SCHEMA on the schema and EXECUTE on the model. If the model declares transitive function dependencies, the creator also needs EXECUTE on those functions. These grants are checked when the endpoint is created and every time it is updated, and a missing grant fails the request with PERMISSION_DENIED. An update also fails if the creator is no longer a workspace member, even when the caller has valid permissions. For this reason, Databricks advises creating endpoints with a long-lived service principal owned by your team, not a personal account that might later be deactivated.

    Checkpoint 6 of 6· Check yourself

    An endpoint was created by an engineer's personal account. The engineer has left and been removed from the workspace, and every configuration update now fails with PERMISSION_DENIED. What is the fix?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Model Serving will install whatever the model imports, so a pyfunc model doesn't need declared dependencies.Why is that wrong?

      The container only gets libraries MLflow captured or that you specified. Native flavors capture their dependencies automatically, but pyfunc models must declare theirs, for example with pip_requirements or extra_pip_requirements.

      Covered in Package dependencies and validate before you deploy

    2. 2.For a Unity Catalog model, entity_name can be just the short model name, such as my-ads-model.Why is that wrong?

      Unity Catalog models need the full three-level name in entity_name, for example catalog.schema.my-ads-model.

      Covered in Create the endpoint: Serving UI, REST API or MLflow Deployments SDK

    3. 3.If the endpoint's creator leaves, an admin can reassign ownership and updates will keep working.Why is that wrong?

      The recorded creator can't be changed after creation and is re-checked on every update. If it is invalid, the endpoint has to be recreated.

      Covered in The endpoint's recorded creator and its Unity Catalog grants

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Log the model or code in the MLflow format, using either native MLflow built-in flavors or pyfunc.”
      ↩︎ From trained model to servable custom model
      “Signatures are necessary for logging models to the Unity Catalog.”
      ↩︎ From trained model to servable custom model
      “For MLflow native flavor models, the necessary package dependencies are automatically captured.”
      ↩︎ Package dependencies and validate before you deploy
      “Model Serving can deploy any Python model or custom code as a production-grade API using CPU or GPU compute resources.”
      ↩︎ Key concept
      “application-level dependencies must be explicitly specified in your MLflow model.”
      ↩︎ Exam trap 1
      “After the model is logged, register it in the Unity Catalog (recommended) or the workspace registry.”
      ↩︎ Checkpoint
      “To include additional requirements beyond what is automatically captured, use extra_pip_requirements.”
      ↩︎ Checkpoint
    2. 2.
      “you can update the requirements by using the MLflow CLI or mlflow.models.model.update_model_requirements() in the MLflow Python API without having to log another model.”
      ↩︎ Package dependencies and validate before you deploy
      “a dependency that is missing, conflicts with another package, or fails to install surfaces here instead of during endpoint deployment.”
      ↩︎ Prediction
    3. 3.
      “MLflow Deployments provides an API for create, update and deletion tasks.”
      ↩︎ Create the endpoint: Serving UI, REST API or MLflow Deployments SDK
      “When you create an endpoint, Databricks records the calling identity as the endpoint's creator.”
      ↩︎ The endpoint's recorded creator and its Unity Catalog grants
      “Use a long-lived service principal owned by your team as the endpoint creator.”
      ↩︎ The endpoint's recorded creator and its Unity Catalog grants
      “provide the full model name including parent catalog and schema such as, catalog.schema.example-model”
      ↩︎ Exam trap 2
      “is used to access Unity Catalog resources on behalf of the endpoint and cannot be changed after creation.”
      ↩︎ Exam trap 3
      “Endpoint names cannot use the databricks- prefix.”
      ↩︎ Checkpoint
      “you must delete the endpoint and recreate it under a service principal that has the required permissions”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Sizing, Scaling and Updating Custom Model Serving Endpoints

    Spotted a mistake, or was something unclear? Tell us.