What you will be able to do
- Choose between the Serving UI, ai_query, the REST API and the MLflow Deployments SDK to query an endpoint
- Authenticate scoring requests with the right kind of token
- Format requests with dataframe_split, dataframe_records, instances or inputs, and read the predictions response
- Update a live endpoint's served version and traffic routes without downtime
1.Four ways to send a scoring request
After a custom model serving endpoint reaches the ready state, you can send it scoring requests in four ways. Each suits a different caller.
| Method | How it works | Typical use |
|---|---|---|
| Serving UI | Query endpoint on the endpoint page; paste JSON, click Send Request (Show Example loads a logged input example) | Fast manual testing |
| SQL function | ai_query invokes the model from SQL | Analytics workflows in SQL |
| REST API | POST /serving-endpoints/{name}/invocations | Web or client applications |
| MLflow Deployments SDK | predict() on a deploy client | Python code |
Checkpoint 1 of 7· Check yourself
An analyst wants to call a served model directly from a SQL query without writing client code. Which querying method fits?
ai_query is the SQL function that invokes model inference directly from SQL, so no client code is needed.
“Invoke model inference directly from SQL using the ai_query SQL function.”Source: docs.databricks.com
https://<databricks-instance>/serving-endpoints/fraud-model/invocations, sent as a POST with a JSON body. Route-optimized endpoints are the exception: they are queried at their own dedicated URL.
2.Authenticating requests
The Serving UI uses your workspace session. Any other caller needs a Databricks API token, whether it uses REST, the MLflow Deployments SDK or a tool like Power BI. Databricks recommends machine-to-machine OAuth tokens in production. For testing and development, it recommends personal access tokens that belong to service principals rather than to workspace users. The MLflow Deployments SDK route also requires MLflow 2.9 or above.
Route-optimized endpoints change these rules. They accept only OAuth tokens, not personal access tokens, and you must query them at a dedicated URL instead of the standard invocations URL.
Checkpoint 2 of 7· Check yourself
A team turns on route optimization for a high-QPS endpoint. Their client keeps sending a service principal's personal access token to /serving-endpoints/{name}/invocations. What is wrong?
Route optimization replaces both the authentication flow and the invocation URL. PATs are not accepted.
“Route-optimized endpoints only accept OAuth tokens (not personal access tokens) and must be queried at a dedicated URL.”Source: docs.databricks.com
Checkpoint 3 of 7· Exam question
An ML engineer registered a model as `main.ml_models.churn_predictor` version 3 in Unity Catalog and now needs to deploy exactly that version behind a Model Serving endpoint. Which configuration correctly identifies the model to serve?
Correct answer: A — A served entity with `entity_name` set to `main.ml_models.churn_predictor` and `entity_version` set to `3`, referencing the Unity Catalog model path and the specific version to deploy.
- A. A served entity's `entity_name` points at the three-level Unity Catalog model path and `entity_version` pins the exact registered version, which is how a serving endpoint is told which model artifact to load.
- B. A `traffic_config` route only distributes request percentages across served models that are already defined; it does not itself identify or register which catalog model version an endpoint loads.
- C. Unity Catalog model versions are referenced by catalog, schema, model name, and version number, not by workspace Model Registry stage aliases like Production, which belong to the legacy registry.
- D. `workload_type` and `scale_to_zero_enabled` configure the compute that runs the endpoint, but neither field identifies which registered model or version the served entity should load.
Sources1
3.Request formats for custom models
Custom models accept two families of input: a Pandas DataFrame serialized as JSON, or tensor input as JSON or Protobuf. You choose the format with a top-level key in the request body.
curl -X POST -u token:$DATABRICKS_API_TOKEN $ENDPOINT_INVOCATION_URL \
-H 'Content-Type: application/json' \
-d '{"dataframe_split": [{
"columns": ["sepal length (cm)", "sepal width (cm)", "petal length (cm)", "petal width (cm)"],
"data": [[5.1, 3.5, 1.4, 0.2], [4.9, 3.0, 1.4, 0.2]]
}]
}'Models that expect tensors, such as TensorFlow or PyTorch models, use instances or inputs. instances is row-oriented. It only works when every input tensor has the same 0-th dimension, because the rows are joined to build each full tensor. inputs is column-oriented, so it can carry named tensors whose instance counts differ. Protobuf requests are not captured in inference tables.
Checkpoint 4 of 7· Match them up
Match each request key to what it sends.
Tap a term, then the definition that fits it.
The two dataframe keys carry tabular input. The two tensor keys differ by orientation: rows for instances, columns for inputs.
“instances is a tensors-based format that accepts tensors in row format.”Source: docs.databricks.com
Sources1
4.Querying from Python and reading the response
From Python, the MLflow Deployments SDK hides the HTTP details. You set the workspace host and token, get a client, and call predict() with the endpoint name and the same JSON body you would send over REST.
client = mlflow.deployments.get_deploy_client("databricks")
response = client.predict(
endpoint="test-model-endpoint",
inputs={"dataframe_split": {
"index": [0, 1],
"columns": ["sepal length (cm)", "sepal width (cm)", "petal length (cm)", "petal width (cm)"],
"data": [[5.1, 3.5, 1.4, 0.2], [4.9, 3.0, 1.4, 0.2]]
}
}
)Every input format gets the same response shape. The model's output is serialized as JSON and wrapped in a predictions key, for example {"predictions": [0, 1, 1, 1, 0]}. Client code can therefore read predictions without caring which input format was sent.
Checkpoint 5 of 7· Check yourself
A client sends a tensor request using the inputs key. What shape will the response have?
DataFrame and tensor requests both get JSON responses with the model output under predictions.
“The response from the endpoint contains the output from your model, serialized with JSON, wrapped in a predictions key.”Source: docs.databricks.com
Sources1
5.Updating a live endpoint without downtime
To promote a new model version, change the endpoint's configuration rather than create a new endpoint. Use PUT /api/2.0/serving-endpoints/{name}/config, call update_endpoint_config in the MLflow Deployments SDK, or click Edit endpoint in the UI. The new config lists the served entities and a traffic_config whose routes give each served model's share of traffic. One endpoint can serve several models or versions at once and split traffic between them.
"traffic_config":
{
"routes": [
{
"served_model_name": "my-ads-model-5",
"traffic_percentage": 100
}
]
}Updates are safe for callers. The old configuration keeps serving predictions until the new one is ready. If the update fails, the existing configuration stays in effect. You can run only one update at a time, but you can cancel an update in progress from the Serving UI. You can change the endpoint name and certain immutable properties, but not most of the rest. Check the endpoint status afterwards to confirm the update was applied.
Checkpoint 6 of 7· Check yourself
You submit a config update that moves an endpoint from version 3 to version 5. While version 5 is still deploying, what serves incoming requests?
Config updates are zero-downtime. The previous configuration handles traffic until the new one is ready, and it stays in effect if the update fails.
“Until the new configuration is ready, the old configuration keeps serving prediction traffic.”Source: docs.databricks.com
Checkpoint 7 of 7· Exam question
A backend service outside Databricks needs to send a scoring request to a serving endpoint named `churn-endpoint` and receive a prediction back over HTTP. Which request correctly queries the endpoint?
Correct answer: A — Send a POST request to `/serving-endpoints/churn-endpoint/invocations` with a personal access token in the `Authorization: Bearer` header and input rows as JSON.
- A. Scoring requests are sent as a POST to the endpoint's `/invocations` path with the input data as a JSON body, authenticated with a bearer token supplied in the Authorization header.
- B. The registered-models API returns MLflow registry metadata rather than predictions, and it does not accept input rows as query parameters or skip authentication for a scoring call.
- C. The invocations endpoint expects a JSON request body describing the input rows, not a CSV file upload, and an authenticated request still requires a bearer token regardless of content type.
- D. Serving endpoint requests authenticate with a Databricks bearer token rather than a username and password, and the invocations path is nested under `/serving-endpoints/<name>/invocations`, not this URL.
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.dataframe_records is the preferred DataFrame format because each row is a self-describing JSON object.Why is that wrong?
dataframe_split is the recommended format. The records format does not guarantee that column order is preserved.
Covered in Request formats for custom models
2.A personal access token from your own user account is the recommended way to authenticate production clients.Why is that wrong?
Databricks recommends machine-to-machine OAuth tokens for production. Personal access tokens, belonging to service principals, are only for testing and development.
Covered in Authenticating requests
Practise it for real
Deploy a registered Unity Catalog model to a real-time endpoint and score it over REST.
1.Create an endpoint with the MLflow Deployments SDK: call get_deploy_client("databricks").create_endpoint with a served_entities entry naming catalog.schema.model and a version, with scale_to_zero_enabled False.
Why: This deploys the registered version into a serving container behind a REST API.
You should see: The endpoint appears on the Serving page as Not Ready while the container builds.
2.Poll the endpoint until state.ready is READY and config_update is NOT_UPDATING.
Why: Requests sent before the served entity reaches DEPLOYMENT_READY cannot be served.
You should see: The served entity's deployment state reads DEPLOYMENT_READY.
3.POST a dataframe_split body to https://<databricks-instance>/serving-endpoints/<endpoint-name>/invocations with a service principal token.
Why: dataframe_split is the recommended format and keeps column order intact.
You should see: A JSON response with your model's output under a predictions key.
4.Send the same rows with client.predict(endpoint=..., inputs={"dataframe_split": ...}).
Why: This confirms the SDK and REST paths reach the same endpoint with the same body.
You should see: Identical predictions to the REST call.
Stuck? Get a nudge
If creation fails with PERMISSION_DENIED, check that the creating identity has USE CATALOG, USE SCHEMA and EXECUTE on the model.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/model-serving/score-custom-model-endpointsOfficial docs
“Invoke model inference directly from SQL using the ai_query SQL function.”
↩︎ Four ways to send a scoring request“Databricks recommends using a personal access token belonging to service principals instead of workspace users.”
↩︎ Authenticating requests“For the MLflow Deployment SDK, MLflow 2.9 or above is required.”
↩︎ Authenticating requests“For custom models, Model Serving supports scoring requests in Pandas DataFrame (JSON) or Tensor input (JSON or Protobuf).”
↩︎ Request formats for custom models“Use MLflow Deployments SDK's predict() function to query the model.”
↩︎ Querying from Python and reading the response“(Recommended)dataframe_split format is a JSON-serialized Pandas DataFrame in the split orientation.”
↩︎ Exam trap 1“Databricks recommends that you use machine-to-machine OAuth tokens for authentication during production.”
↩︎ Exam trap 2“Route-optimized endpoints only accept OAuth tokens (not personal access tokens) and must be queried at a dedicated URL.”
↩︎ Checkpoint“This format does not guarantee the preservation of column ordering, and the split format is preferred over the records format.”
↩︎ Prediction“instances is a tensors-based format that accepts tensors in row format.”
↩︎ Checkpoint“The response from the endpoint contains the output from your model, serialized with JSON, wrapped in a predictions key.”
↩︎ Checkpoint - 2.
“The easiest and fastest way to test and send scoring requests to your served model is to use the Serving UI.”
↩︎ Four ways to send a scoring request - 3.https://docs.databricks.com/aws/en/machine-learning/model-serving/create-manage-serving-endpointsOfficial docs
“You can serve multiple models or model versions from a single endpoint and control the traffic split between them.”
↩︎ Updating a live endpoint without downtime“While there is an update in progress, another update cannot be made.”
↩︎ Updating a live endpoint without downtime“Until the new configuration is ready, the old configuration keeps serving prediction traffic.”
↩︎ Checkpoint