What you will be able to do
- Explain what Foundation Model APIs give an LLM application compared with deploying and maintaining your own model
- Choose between pay-per-token, priority pay-per-token, provisioned throughput (on-demand or reserved), and AI Functions for a given workload
- List the requirements for using Foundation Model APIs
- Create a provisioned throughput endpoint from system.ai, through the UI or the REST API, and size it using the model optimization information API
Key concept
Foundation Model APIs — A Model Serving capability where Databricks hosts the foundation models and exposes them as serving endpoints. Your LLM application calls the endpoint and never deploys or operates the model. What you do choose is the serving mode, and that choice sets the cost, latency consistency, and capacity guarantees.
1.What Foundation Model APIs give an LLM application
Every LLM application needs a model it can call. On Databricks the shortest path is Foundation Model APIs, a feature of Model Serving that lets you "access and query state-of-the-art open models from a serving endpoint." Databricks hosts the models, so your work is the application code and not the model infrastructure.
The documentation lists what this is for: checking whether a project is viable before you invest more, building a quick proof of concept, building a RAG chatbot by pairing a foundation model with a vector index, replacing proprietary models with open ones, comparing LLMs or swapping in a better one, and running production apps on "a scalable, SLA-backed LLM serving solution." The models are also a Databricks Designated Service, which means Databricks Geos manage data residency when customer content is processed.
Checkpoint 1 of 5· Check yourself
Compared with deploying your own model, what do Foundation Model APIs let a team skip?
Databricks hosts the models, so the team builds the application without maintaining a model deployment. They still write application code, choose a serving mode, and authenticate.
“These models are hosted by Databricks and you can quickly and easily build applications that use them without maintaining your own model deployment.”Source: docs.databricks.com
Sources1
2.Picking a serving mode
The model is the same in every mode. What changes is how you pay for capacity and what performance you get. Pay-per-token is the starting point: Databricks preconfigures endpoints for the pay-per-token models, and you find them by clicking AI Gateway in the sidebar, where they are listed as system-provided model services.
For latency-sensitive production traffic on pay-per-token there is priority mode. You set the service_tier request parameter to "priority", and Databricks admits the request ahead of standard best-effort pay-per-token traffic. You opt in per request, there is no commitment, and you pay a higher per-token rate.
Provisioned throughput gives you dedicated capacity. It is the recommended mode for production, and it is required when you serve fine-tuned or fully custom weights. It comes in two options: on-demand, with no term, and reserved, with a fixed 1- or 3-month term at a lower per-unit rate. Provisioned throughput endpoints are also available with compliance certifications such as HIPAA. The last mode, AI Functions optimized models, is meant for batch inference.
| Mode | Commitment / billing | Recommended for |
|---|---|---|
| Pay-per-token | No commitment; billed per token | Getting started; preconfigured endpoints are already in the workspace |
| Priority pay-per-token (service_tier set to "priority") | Opt-in per request, no commitment, higher per-token rate | Latency-sensitive production workloads that need consistent performance under load |
| Provisioned throughput: on-demand | Dedicated capacity, no term; change or remove it at any time | Production workloads that need high throughput, performance guarantees, fine-tuned models, or additional security |
| Provisioned throughput: reserved | Dedicated capacity reserved for a fixed 1- or 3-month term | The same production needs, with capacity committed for a fixed term |
| AI Functions optimized models | Run through AI Functions | Batch inference workloads |
Checkpoint 2 of 5· Check yourself
A real-time support assistant runs on pay-per-token and has inconsistent latency at peak. The team does not want dedicated capacity or a term commitment. What should they do?
Priority pay-per-token is opt-in per request and needs no commitment. It admits requests ahead of best-effort traffic so latency stays more consistent. Reserved provisioned throughput means committing to a term, and AI Functions mode is for batch work.
“Priority mode is opt-in per request, requires no commitment, and is billed at a higher per-token rate.”Source: docs.databricks.com
Sources1
3.Requirements and model task types
Before an application can call Foundation Model APIs, three things must be in place. You need a Databricks API token to authenticate requests. You need serverless compute if you use provisioned throughput models. Your workspace must be in a supported region, and the pay-per-token and provisioned throughput modes each have their own list of regions.
The models fall into task types, and the task type tells you which model a feature needs. General purpose (chat) models handle multi-turn conversation, for example virtual assistants and support bots. Embedding models turn data into vectors for semantic search and RAG; the Databricks-hosted ones are databricks-qwen3-embedding-0-6b, databricks-gte-large-en and databricks-bge-large-en. Vision models analyse images and documents, and reasoning models cover code generation and agent orchestration. A RAG application usually needs two of these at once: an embedding model for retrieval and a chat model for generation.
Checkpoint 3 of 5· Match them up
Match each task type to the use case the documentation recommends it for
Tap a term, then the definition that fits it.
Each task type maps to its recommended use cases in the foundation model types table.
“Recommended for applications where semantic understanding, similarity comparison, and efficient retrieval or clustering of complex data are essential”Source: docs.databricks.com
Serverless compute. The requirements list it as "Serverless compute (for provisioned throughput models)." The API token and a supported region apply to both modes.
Sources1
4.Creating a provisioned throughput endpoint
Databricks recommends serving the foundation models that come pre-installed in Unity Catalog, in the system catalog under the ai schema (system.ai). In the UI you open system.ai in Catalog Explorer, click the model, then click Serve this model. That opens the Create serving endpoint page. In the Up to dropdown you set the maximum tokens per second. The endpoint scales automatically, and Modify shows the minimum it can scale down to. One restriction: for Meta Llama models you must choose an Instruct version, because base versions cannot be deployed from Unity Catalog.
Checkpoint 4 of 5· Put it in order
Put the UI steps for serving a foundation model from Unity Catalog in order
- 1.Navigate to system.ai in Catalog Explorer
- 2.Configure the endpoint on the Create serving endpoint page
- 3.Click on the name of the model to deploy
- 4.On the model page, click the Serve this model button
You start at the system.ai schema, choose the model, click Serve this model, and land on the Create serving endpoint page.
“On the model page, click the Serve this model button.”Source: docs.databricks.com
With the REST API, the request body depends on how the model measures capacity: in throughput bands or in model units. In both cases you first call the model optimization information API, GET api/2.0/serving-endpoints/get-model-optimization-info/{registered_model_name}/{version}. Its response says whether the model is optimizable and gives the chunk size, which is the increment you provision in.
{
"optimizable": true,
"model_type": "llama",
"throughput_chunk_size": 980
}| Capacity measure | Chunk size field | Request fields | POST path |
|---|---|---|---|
| Throughput bands | throughput_chunk_size | min_provisioned_throughput and max_provisioned_throughput | /api/2.0/serving-endpoints |
| Model units | model_unit_chunk_size | provisioned_model_units | /api/2.0/serving-endpoints/pt |
# Send the POST request to create the serving endpoint
data = {
"name": endpoint_name,
"config": {
"served_entities": [
{
"entity_name": model_name,
"entity_version": model_version,
"provisioned_model_units": desired_model_units,
}
]
},
}
response = requests.post(
url=f"{API_ROOT}/api/2.0/serving-endpoints/pt", json=data, headers=headers
)Checkpoint 5 of 5· Fill the gap
For a model that measures capacity in model units, which key of the optimization info response gives the increment?
chunk_size = optimizable_info[' ? ']
# Desired provisioned throughput
desired_model_units = 2 * chunk_sizeModel-unit models return model_unit_chunk_size, and throughput-band models return throughput_chunk_size. provisioned_model_units is the field you set in the request body, not a field in the response.
Source: docs.databricks.comSources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Provisioned throughput is the easiest way to start using Foundation Model APIs.Why is that wrong?
Pay-per-token is the recommended starting point, and its endpoints are already preconfigured in the workspace. Provisioned throughput is the recommendation for production workloads that need guarantees.
Covered in Picking a serving mode
2.Every provisioned throughput endpoint is created by POSTing min/max throughput to /api/2.0/serving-endpoints.Why is that wrong?
Only throughput-band models use min_provisioned_throughput and max_provisioned_throughput. Model-unit models use provisioned_model_units and POST to the /pt path.
Covered in Creating a provisioned throughput endpoint
3.Any Meta Llama variant in system.ai can be served on provisioned throughput.Why is that wrong?
Only Instruct versions can be deployed from Unity Catalog. Base versions are not supported.
Covered in Creating a provisioned throughput endpoint
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“These models are hosted by Databricks and you can quickly and easily build applications that use them without maintaining your own model deployment.”
↩︎ What Foundation Model APIs give an LLM application“This mode is recommended for all production workloads, especially those that require high throughput, performance guarantees, fine-tuned models, or have additional security requirements.”
↩︎ Picking a serving mode“AI Functions optimized models: This mode is recommended for batch inference workloads.”
↩︎ Picking a serving mode“Reserved provisioned throughput: reserve dedicated capacity for a fixed 1- or 3-month term.”
↩︎ Picking a serving mode“Serverless compute (for provisioned throughput models).”
↩︎ Requirements and model task types“Databricks API token to authenticate endpoint requests.”
↩︎ Requirements and model task types“These models are hosted by Databricks and you can quickly and easily build applications that use them without maintaining your own model deployment.”
↩︎ Key concept“This is the easiest way to start accessing foundation models on Databricks and is recommended for beginning your journey with Foundation Model APIs.”
↩︎ Exam trap 1“Use a foundation model, along with a vector index, to build a chatbot using retrieval augmented generation (RAG).”
↩︎ Prediction“Priority mode is opt-in per request, requires no commitment, and is billed at a higher per-token rate.”
↩︎ Checkpoint - 2.
“Databricks recommends using the foundation models that are pre-installed in Unity Catalog.”
↩︎ Creating a provisioned throughput endpoint“To identify the suitable range for your needs, Databricks recommends using the model optimization information API within the platform.”
↩︎ Creating a provisioned throughput endpoint“Provisioned throughput endpoints automatically scale, so you can select Modify to view the minimum tokens per second your endpoint can scale down to.”
↩︎ Creating a provisioned throughput endpoint“For these models, you specify the provisioned_model_units field in your request and send the POST request to the /api/2.0/serving-endpoints/pt endpoint.”
↩︎ Exam trap 2“Base versions of the Meta Llama models are not supported for deployment from Unity Catalog.”
↩︎ Exam trap 3“On the model page, click the Serve this model button.”
↩︎ Checkpoint
Also cited
- https://docs.databricks.com/aws/en/machine-learning/model-serving/score-foundation-modelsOfficial docs
“Recommended for applications where semantic understanding, similarity comparison, and efficient retrieval or clustering of complex data are essential”
↩︎ Checkpoint