CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 4 · Lesson 47/48

    Query a Real-Time Model Serving Endpoint

    Deploy and query a model for realtime inference

    10 min read
    2.08% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Choose between the Serving UI, ai_query, the REST API and the MLflow Deployments SDK to query an endpoint
    • Authenticate scoring requests with the right kind of token
    • Format requests with dataframe_split, dataframe_records, instances or inputs, and read the predictions response
    • Update a live endpoint's served version and traffic routes without downtime

    1.Four ways to send a scoring request

    After a custom model serving endpoint reaches the ready state, you can send it scoring requests in four ways. Each suits a different caller.

    Options for querying a custom model serving endpoint
    MethodHow it worksTypical use
    Serving UIQuery endpoint on the endpoint page; paste JSON, click Send Request (Show Example loads a logged input example)Fast manual testing
    SQL functionai_query invokes the model from SQLAnalytics workflows in SQL
    REST APIPOST /serving-endpoints/{name}/invocationsWeb or client applications
    MLflow Deployments SDKpredict() on a deploy clientPython code

    Checkpoint 1 of 7· Check yourself

    An analyst wants to call a served model directly from a SQL query without writing client code. Which querying method fits?

    Sources12

    2.Authenticating requests

    The Serving UI uses your workspace session. Any other caller needs a Databricks API token, whether it uses REST, the MLflow Deployments SDK or a tool like Power BI. Databricks recommends machine-to-machine OAuth tokens in production. For testing and development, it recommends personal access tokens that belong to service principals rather than to workspace users. The MLflow Deployments SDK route also requires MLflow 2.9 or above.

    Route-optimized endpoints change these rules. They accept only OAuth tokens, not personal access tokens, and you must query them at a dedicated URL instead of the standard invocations URL.

    Checkpoint 2 of 7· Check yourself

    A team turns on route optimization for a high-QPS endpoint. Their client keeps sending a service principal's personal access token to /serving-endpoints/{name}/invocations. What is wrong?

    Checkpoint 3 of 7· Exam question

    An ML engineer registered a model as `main.ml_models.churn_predictor` version 3 in Unity Catalog and now needs to deploy exactly that version behind a Model Serving endpoint. Which configuration correctly identifies the model to serve?

    Sources1

    3.Request formats for custom models

    Custom models accept two families of input: a Pandas DataFrame serialized as JSON, or tensor input as JSON or Protobuf. You choose the format with a top-level key in the request body.

    Scoring with the recommended dataframe_split format over RESTbash
    curl -X POST -u token:$DATABRICKS_API_TOKEN $ENDPOINT_INVOCATION_URL \
      -H 'Content-Type: application/json' \
      -d '{"dataframe_split": [{
        "columns": ["sepal length (cm)", "sepal width (cm)", "petal length (cm)", "petal width (cm)"],
        "data": [[5.1, 3.5, 1.4, 0.2], [4.9, 3.0, 1.4, 0.2]]
        }]
      }'

    Models that expect tensors, such as TensorFlow or PyTorch models, use instances or inputs. instances is row-oriented. It only works when every input tensor has the same 0-th dimension, because the rows are joined to build each full tensor. inputs is column-oriented, so it can carry named tensors whose instance counts differ. Protobuf requests are not captured in inference tables.

    Checkpoint 4 of 7· Match them up

    Match each request key to what it sends.

    Tap a term, then the definition that fits it.

    Sources1

    4.Querying from Python and reading the response

    From Python, the MLflow Deployments SDK hides the HTTP details. You set the workspace host and token, get a client, and call predict() with the endpoint name and the same JSON body you would send over REST.

    Querying an endpoint with the MLflow Deployments SDK predict() APIpython
    client = mlflow.deployments.get_deploy_client("databricks")
    
    response = client.predict(
                endpoint="test-model-endpoint",
                inputs={"dataframe_split": {
                        "index": [0, 1],
                        "columns": ["sepal length (cm)", "sepal width (cm)", "petal length (cm)", "petal width (cm)"],
                        "data": [[5.1, 3.5, 1.4, 0.2], [4.9, 3.0, 1.4, 0.2]]
                        }
                    }
              )

    Every input format gets the same response shape. The model's output is serialized as JSON and wrapped in a predictions key, for example {"predictions": [0, 1, 1, 1, 0]}. Client code can therefore read predictions without caring which input format was sent.

    Checkpoint 5 of 7· Check yourself

    A client sends a tensor request using the inputs key. What shape will the response have?

    Sources1

    5.Updating a live endpoint without downtime

    To promote a new model version, change the endpoint's configuration rather than create a new endpoint. Use PUT /api/2.0/serving-endpoints/{name}/config, call update_endpoint_config in the MLflow Deployments SDK, or click Edit endpoint in the UI. The new config lists the served entities and a traffic_config whose routes give each served model's share of traffic. One endpoint can serve several models or versions at once and split traffic between them.

    The traffic_config portion of a config update, sending all traffic to version 5json
        "traffic_config":
        {
          "routes": [
            {
              "served_model_name": "my-ads-model-5",
              "traffic_percentage": 100
            }
          ]
        }

    Updates are safe for callers. The old configuration keeps serving predictions until the new one is ready. If the update fails, the existing configuration stays in effect. You can run only one update at a time, but you can cancel an update in progress from the Serving UI. You can change the endpoint name and certain immutable properties, but not most of the rest. Check the endpoint status afterwards to confirm the update was applied.

    Checkpoint 6 of 7· Check yourself

    You submit a config update that moves an endpoint from version 3 to version 5. While version 5 is still deploying, what serves incoming requests?

    Checkpoint 7 of 7· Exam question

    A backend service outside Databricks needs to send a scoring request to a serving endpoint named `churn-endpoint` and receive a prediction back over HTTP. Which request correctly queries the endpoint?

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.dataframe_records is the preferred DataFrame format because each row is a self-describing JSON object.Why is that wrong?

      dataframe_split is the recommended format. The records format does not guarantee that column order is preserved.

      Covered in Request formats for custom models

    2. 2.A personal access token from your own user account is the recommended way to authenticate production clients.Why is that wrong?

      Databricks recommends machine-to-machine OAuth tokens for production. Personal access tokens, belonging to service principals, are only for testing and development.

      Covered in Authenticating requests

    Practise it for real

    Deploy a registered Unity Catalog model to a real-time endpoint and score it over REST.

    1. 1.Create an endpoint with the MLflow Deployments SDK: call get_deploy_client("databricks").create_endpoint with a served_entities entry naming catalog.schema.model and a version, with scale_to_zero_enabled False.

      Why: This deploys the registered version into a serving container behind a REST API.

      You should see: The endpoint appears on the Serving page as Not Ready while the container builds.

    2. 2.Poll the endpoint until state.ready is READY and config_update is NOT_UPDATING.

      Why: Requests sent before the served entity reaches DEPLOYMENT_READY cannot be served.

      You should see: The served entity's deployment state reads DEPLOYMENT_READY.

    3. 3.POST a dataframe_split body to https://<databricks-instance>/serving-endpoints/<endpoint-name>/invocations with a service principal token.

      Why: dataframe_split is the recommended format and keeps column order intact.

      You should see: A JSON response with your model's output under a predictions key.

    4. 4.Send the same rows with client.predict(endpoint=..., inputs={"dataframe_split": ...}).

      Why: This confirms the SDK and REST paths reach the same endpoint with the same body.

      You should see: Identical predictions to the REST call.

    Stuck? Get a nudge

    If creation fails with PERMISSION_DENIED, check that the creating identity has USE CATALOG, USE SCHEMA and EXECUTE on the model.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Invoke model inference directly from SQL using the ai_query SQL function.”
      ↩︎ Four ways to send a scoring request
      “Databricks recommends using a personal access token belonging to service principals instead of workspace users.”
      ↩︎ Authenticating requests
      “For the MLflow Deployment SDK, MLflow 2.9 or above is required.”
      ↩︎ Authenticating requests
      “For custom models, Model Serving supports scoring requests in Pandas DataFrame (JSON) or Tensor input (JSON or Protobuf).”
      ↩︎ Request formats for custom models
      “Use MLflow Deployments SDK's predict() function to query the model.”
      ↩︎ Querying from Python and reading the response
      “(Recommended)dataframe_split format is a JSON-serialized Pandas DataFrame in the split orientation.”
      ↩︎ Exam trap 1
      “Databricks recommends that you use machine-to-machine OAuth tokens for authentication during production.”
      ↩︎ Exam trap 2
      “Route-optimized endpoints only accept OAuth tokens (not personal access tokens) and must be queried at a dedicated URL.”
      ↩︎ Checkpoint
      “This format does not guarantee the preservation of column ordering, and the split format is preferred over the records format.”
      ↩︎ Prediction
      “instances is a tensors-based format that accepts tensors in row format.”
      ↩︎ Checkpoint
      “The response from the endpoint contains the output from your model, serialized with JSON, wrapped in a predictions key.”
      ↩︎ Checkpoint
    2. 2.
      “The easiest and fastest way to test and send scoring requests to your served model is to use the Serving UI.”
      ↩︎ Four ways to send a scoring request
    3. 3.
      “You can serve multiple models or model versions from a single endpoint and control the traffic split between them.”
      ↩︎ Updating a live endpoint without downtime
      “While there is an update in progress, another update cannot be made.”
      ↩︎ Updating a live endpoint without downtime
      “Until the new configuration is ready, the old configuration keeps serving prediction traffic.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.