What you will be able to do
- Choose a query method (OpenAI client, AI Functions, Serving UI, REST API, MLflow Deployments SDK, Databricks Python SDK) for an application
- Apply the recommended authentication for development and for production
- Write a chat request with the OpenAI-compatible client and read the response
- Run batch inference over a foundation model from SQL with ai_query
- Route an agent's LLM calls through a Unity Gateway endpoint
1.Ways an application can call a foundation model
After you pick a serving mode, the application needs a way to send requests. Model Serving uses one OpenAI-compatible API for Databricks-hosted Foundation Model APIs and for external models. That lets the same calling code work whichever provider is behind the endpoint. All of these requests go through Unity Gateway, where you can apply rate limits, budgets and guardrails.
If you just want to call an LLM, the docs suggest model services in Unity Gateway. These give you governed access to Databricks-hosted models without configuring a serving endpoint. If you serve the model on an endpoint you host yourself, the calling code is the same except for one value: you replace the model service name (for example system.ai.claude-sonnet-4-5) with your endpoint name.
| Method | How it works | Setup |
|---|---|---|
| OpenAI client | Pass the endpoint (or model service) name as the model input; supports chat, embeddings and completions | pip install -U databricks-openai |
| AI Functions | Call model inference from SQL with ai_query | None beyond SQL access |
| Serving UI | Query endpoint on the Serving endpoint page; Show Example loads a logged input example | None |
| REST API | POST /serving-endpoints/{name}/invocations | Databricks API token |
| MLflow Deployments SDK | Call predict() on a deployments client | pip install mlflow |
| Databricks Python SDK | A layer over the REST API that handles authentication | Preinstalled on Databricks Runtime 13.3 LTS or above |
Checkpoint 1 of 6· Check yourself
A developer is building a long-running chat application on Foundation Model APIs. Which interface do the docs recommend for extended interactions?
Because the APIs are OpenAI-compatible, Databricks recommends the OpenAI client for extended interactions. The UI is the recommendation for trying the feature out.
“Databricks recommends using the OpenAI client SDK or API for extended interactions and the UI for trying out the feature.”Source: docs.databricks.com
Sources1
2.Authenticating the application
The OpenAI client, the REST API and the MLflow Deployments SDK all need a Databricks API token to send scoring requests. Which token depends on the stage. For testing and development, use a personal access token that belongs to a service principal rather than to a workspace user. In production, switch to machine-to-machine OAuth tokens.
The databricks-openai package means you don't have to wire up credentials by hand: it "provides an OpenAI client with authorization automatically configured." In a notebook you install it and then run dbutils.library.restartPython(). You only use the plain openai client, with your token as api_key and the workspace's /ai-gateway/mlflow/v1 URL as base_url, when you query foundation models from outside the workspace.
Checkpoint 2 of 6· Check yourself
A data scientist is testing an endpoint from a dev notebook. Which token do the docs recommend?
For testing and development, the docs recommend service-principal PATs instead of tokens tied to workspace users.
“For testing and development, Databricks recommends using a personal access token belonging to service principals instead of workspace users.”Source: docs.databricks.com
Checkpoint 3 of 6· Exam question
A developer wants to minimize code changes while switching their application between a Databricks-hosted Foundation Model APIs endpoint and a third-party hosted LLM, and also wants to reuse existing tooling built around a widely adopted client library. Which client interface does Databricks recommend for querying Foundation Model APIs endpoints in this scenario?
Correct answer: A — The OpenAI-compatible client SDK, since Foundation Model APIs endpoints accept requests in the OpenAI chat completions format
- A. The OpenAI-compatible client SDK is correct because Databricks recommends it for Foundation Model APIs, letting requests use the familiar OpenAI chat completions format and making it easy to swap between Databricks-hosted and other OpenAI-compatible endpoints.
- B. The Feature Engineering client is used to read and write feature tables for machine learning features, not to send inference requests to an LLM serving endpoint.
- C. Delta Sharing distributes data and tabular assets across organizations and does not provide a request/response interface for querying a live model serving endpoint.
- D. MLlib's prediction interfaces target classic distributed Spark ML models and are not built for the chat/completions request format used by foundation model serving endpoints.
Sources1
3.Sending a chat request and reading the response
Inside the workspace, a chat call is a standard OpenAI chat.completions.create call. Only the model value is Databricks-specific.
from databricks_openai import DatabricksOpenAI
client = DatabricksOpenAI()
response = client.chat.completions.create(
model="system.ai.claude-sonnet-4-5",
messages=[
{
"role": "system",
"content": "You are a helpful assistant."
},
{
"role": "user",
"content": "What is a mixture of experts model?",
}
],
max_tokens=256
)A REST call carries the same fields in a JSON body: messages, max_tokens and temperature. The response follows the OpenAI chat completion shape, including a usage object that counts prompt and completion tokens. On pay-per-token, that token count is what you are billed for.
{
"model": "databricks-claude-sonnet-4-5",
"choices": [
{
"message": {},
"index": 0,
"finish_reason": null
}
],
"usage": {
"prompt_tokens": 7,
"completion_tokens": 74,
"total_tokens": 81
},
"object": "chat.completion",
"id": null,
"created": 1698824353
}Checkpoint 4 of 6· Check yourself
Your app queries the model service system.ai.claude-sonnet-4-5. You move it to a model serving endpoint that you host. What changes in the client call?
The examples work for both. Only the name in the model input changes.
“If you use model serving endpoints instead of model services, replace the model service name with an endpoint name.”Source: docs.databricks.com
Sources2
4.Batch inference from SQL with ai_query
Not every LLM application serves one user at a time. To transform rows in a table, you can call the same foundation models from SQL with the built-in ai_query function. This is the AI Functions path, and the serving mode recommended for batch inference is AI Functions optimized models. Note that the function is in Public Preview and its definition might change.
Checkpoint 5 of 6· Fill the gap
Which SQL function sends this prompt to a foundation model?
SELECT ? (
"system.ai.claude-sonnet-4-5",
"Can you explain AI in ten words?"
)ai_query is the built-in SQL function that runs model inference directly from SQL.
Source: docs.databricks.com5.Serving an agent app that calls Foundation Model APIs
A whole LLM application, such as an agent with a chat UI, can be deployed on Databricks Apps. Apps "gives you full control over the agent code, server configuration, and deployment workflow." The app still gets its model from a foundation model endpoint. The recommended way to wire this up is to send every LLM call through Unity Gateway: in the agent code, pass the gateway endpoint name as model and set use_ai_gateway=True on the Databricks LLM client (AsyncDatabricksOpenAI or ChatDatabricks in LangChain). The client then handles authentication. Because the gateway sits in the request path, you can swap models, attribute cost per app, and inspect or replay traffic without changing agent code or rotating provider credentials.
Checkpoint 6 of 6· Check yourself
An agent on Databricks Apps uses ChatDatabricks. What makes its LLM calls go through Unity Gateway?
Two settings route the calls: the gateway endpoint name as the model argument, and use_ai_gateway=True on the client.
“In your agent code, pass the Unity Gateway endpoint name as the model argument and set use_ai_gateway=True on the Databricks LLM client.”Source: docs.databricks.com
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A developer's personal access token is the right credential for an app querying Foundation Model APIs in production.Why is that wrong?
Production should use machine-to-machine OAuth tokens. Even for development, the recommendation is service-principal PATs, not tokens tied to workspace users.
Covered in Authenticating the application
2.DatabricksOpenAI works the same way from outside the workspace.Why is that wrong?
To query from outside the workspace you use the plain OpenAI client, configured with your token and the workspace URL.
Covered in Authenticating the application
Practise it for real
Query a pay-per-token foundation model from a notebook, first with the OpenAI-compatible client and then from SQL
1.In a notebook, run pip install -U databricks-openai, then dbutils.library.restartPython()
Why: The package provides an OpenAI client with authorization configured for Databricks
You should see: The package installs and the Python process restarts
2.Run the DatabricksOpenAI chat.completions.create example with model="system.ai.claude-sonnet-4-5"
Why: This is the recommended OpenAI-compatible path for calling a model service
You should see: A chat completion object whose usage shows prompt, completion and total tokens
3.Change the model argument to a different general purpose model from the supported list and run it again
Why: Shows that the shared API lets you compare or swap models without changing any other code
You should see: A response in the same format from the other model
4.In a SQL cell, run SELECT ai_query("system.ai.claude-sonnet-4-5", "Can you explain AI in ten words?")
Why: ai_query is the AI Functions path for model inference from SQL, used for batch workloads
You should see: A single text response returned as the query result
Stuck? Get a nudge
If the model name is rejected, check whether you are calling a model service (system.ai.*) or a serving endpoint, and use the matching name.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/model-serving/score-foundation-modelsOfficial docs
“routed through Unity Gateway, which lets you apply rate limits, budgets, and guardrails to control cost and access.”
↩︎ Ways an application can call a foundation model“Model services give you governed access to foundation models served natively by Databricks through a single interface, without configuring a serving endpoint.”
↩︎ Ways an application can call a foundation model“This package provides an OpenAI client with authorization automatically configured to query AI models.”
↩︎ Authenticating the application“Invoke model inference directly from SQL using the ai_query SQL function.”
↩︎ Batch inference from SQL with ai_query“As a security best practice for production scenarios, Databricks recommends that you use machine-to-machine OAuth tokens for authentication during production.”
↩︎ Exam trap 1“As a security best practice for production scenarios, Databricks recommends that you use machine-to-machine OAuth tokens for authentication during production.”
↩︎ Prediction“For testing and development, Databricks recommends using a personal access token belonging to service principals instead of workspace users.”
↩︎ Checkpoint - 2.
“If you use model serving endpoints instead of model services, replace the model service name with an endpoint name.”
↩︎ Sending a chat request and reading the response“This function is in Public Preview and the definition might change.”
↩︎ Batch inference from SQL with ai_query“To query foundation models outside of your workspace, you must use the OpenAI client directly.”
↩︎ Exam trap 2 - 3.
“you can centralize permissions, attribute cost per app, swap models, and inspect or replay traffic without modifying agent code or rotating provider credentials.”
↩︎ Serving an agent app that calls Foundation Model APIs“Databricks Apps gives you full control over the agent code, server configuration, and deployment workflow.”
↩︎ Serving an agent app that calls Foundation Model APIs“In your agent code, pass the Unity Gateway endpoint name as the model argument and set use_ai_gateway=True on the Databricks LLM client.”
↩︎ Checkpoint
Also cited
“Databricks recommends using the OpenAI client SDK or API for extended interactions and the UI for trying out the feature.”
↩︎ Checkpoint