What you will be able to do
- Pick pay-per-token, priority mode, provisioned throughput or AI Functions based on an application's production needs
- Use rate limits and regional availability to confirm a candidate model can serve the expected load
- Confirm a model choice by comparing candidates on quality, cost and latency metrics
1.Match the serving mode to the workload
Choosing an LLM involves more than the model's capabilities. It also depends on how the model will be served. Databricks Foundation Model APIs offer several modes, and each one suits a different kind of application. When you pick a model, you also need to pick a mode that model supports.
| Mode | Fits applications that… | Notes from the docs |
|---|---|---|
| Pay-per-token | are just starting out or are proofs of concept | Preconfigured endpoints in the workspace; recommended for getting started |
| Priority pay-per-token (priority mode) | are latency-sensitive and real-time | Opt-in per request, no commitment, higher per-token rate |
| Provisioned throughput | need high throughput, performance guarantees, fine-tuned models or extra security | On-demand (no term) or reserved (fixed 1- or 3-month term); HIPAA available |
| AI Functions optimized models | run batch inference | Recommended for batch inference workloads |
Provisioned throughput also matters when the best model is one you customised yourself. It supports "Fine-tuned variants of base models, such as models that are fine-tuned on proprietary data," and base models registered from Hugging Face or another external source, an approach that "works with any fine-tuned variant of the supported models." If your application needs a model that isn't hosted by Databricks at all, external models let you reach third-party providers such as OpenAI, Anthropic, Cohere, Amazon Bedrock and Google Cloud Vertex AI through one endpoint interface.
Checkpoint 1 of 5· Match them up
Match each application to the Foundation Model APIs mode recommended for it
Tap a term, then the definition that fits it.
Batch goes to AI Functions, latency-sensitive opt-in traffic goes to priority mode, fine-tuned or guaranteed production workloads go to provisioned throughput, and getting started goes to pay-per-token.
“This mode is recommended for all production workloads, especially those that require high throughput, performance guarantees, fine-tuned models, or have additional security requirements.”Source: docs.databricks.com
Checkpoint 2 of 5· Exam question
A legal team wants to summarize full contracts, some running 50 to 100 pages, in a single prompt without splitting the document into chunks. Which model attribute should be prioritized when selecting an LLM for this application?
Correct answer: A — A large context window that can accept the entire contract in a single prompt, avoiding the need to split the document into chunks
- A large context window that can accept the entire contract in a single prompt, avoiding the need to split the document into chunks. Prioritizing a large context window is correct because the application's defining constraint is fitting long documents into one prompt, and an insufficient context window would force chunking and risk losing cross-section context in the contract.
- The highest reported benchmark score on coding tasks, since strong coding performance typically transfers to accurate long-document summarization. This is incorrect because coding benchmarks measure a different capability and do not indicate whether a model can accept or reason over very long input documents.
- The lowest cost per token, since summarization quality depends primarily on minimizing inference spend rather than input capacity. This is incorrect because a model with an inadequate context window is unusable for this task regardless of price, so input capacity is the binding constraint rather than cost.
- Fine-tuning for real-time voice transcription, since transcription-focused training improves comprehension of long written contracts. This is incorrect because voice transcription tuning targets audio-to-text conversion and has no bearing on a model's ability to process long written legal text.
2.Check that the model can carry the load
A model that fits the task can still fail the application if it can't handle the traffic. Pay-per-token endpoints have per-model limits on input tokens per minute (ITPM), output tokens per minute (OTPM) and queries per hour (QPH), and those limits differ a lot between models. The figures below are for Enterprise tier workspaces.
| Model | ITPM limit | OTPM limit | QPH limit |
|---|---|---|---|
| GPT OSS 20B | 1,000,000 | 100,000 | 360,000 |
| GPT OSS 120B | 1,000,000 | 100,000 | 360,000 |
| Claude Opus 5 | 1,000,000 | 100,000 | 360,000 |
| Claude Opus 4.8 | 200,000 | 20,000 | 360,000 |
| GLM 5.3 | 2,000,000 | 40,000 | 7,200 |
| DeepSeek V4 Pro (0813) | 200,000 | 4,000 | 7,200 |
Read the table against your application's traffic. GLM 5.3 accepts a very high input-token rate but allows only 7,200 queries per hour, so an app that sends many short requests will hit the QPH limit long before the token limit. The limits page also links model size to workload: "Smaller models for high-volume tasks: Use models like GPT OSS 20B for tasks that require higher throughput," and reserve GPT OSS 120B for tasks that require maximum capability. Some models also cap request size; for Gemini 2.5 Pro and Gemini 2.5 Flash, "Requests must be smaller than 200K input tokens or 400KB."
If a model is otherwise the right fit but its limits are too low, you don't have to switch models. The page says: "For models that support provisioned throughput, create a provisioned throughput endpoint if the pay-per-token limits don't meet your use case requirements."
Finally, confirm that the model is actually available to you. Model assets in system.ai are listed globally, and "Seeing a model there does not mean it is available in your workspace." Availability depends on region, cross-geo settings and the model itself, so check the Unity Gateway UI and your permissions before you commit.
Checkpoint 3 of 5· Check yourself
Evaluation shows Claude Opus 4.8 is the best-quality fit, but projected traffic exceeds its pay-per-token limits. The model supports provisioned throughput. What do the docs recommend?
When pay-per-token limits don't meet the use case and the model supports provisioned throughput, the docs say to create a provisioned throughput endpoint. Priority mode changes admission order, not the per-model limits.
“create a provisioned throughput endpoint if the pay-per-token limits don't meet your use case requirements.”Source: docs.databricks.com
3.Confirm the choice by comparing candidates
Model descriptions and limit tables give you a shortlist, not a final answer. Every entry on the supported-models page carries the same warning: "Databricks recommends customers evaluate their models to ensure they're operating as intended." The Foundation Model APIs make it cheap to do that. One listed use is to "Efficiently compare LLMs to see the best candidate for your use case, or swap a production model with a better performing one." Another is to replace proprietary models with open alternatives "to optimize for cost and performance."
A comparison needs the same metrics for every candidate. Quality is one of them. System performance is another: "Overall latency and token consumption are examples of chain performance metrics," and both can be measured deterministically from the application's outputs. Token counts matter twice: they drive cost, and they show whether a candidate stays within its ITPM and OTPM limits. The limits page shows how to read them from each response.
Checkpoint 4 of 5· Fill the gap
Complete the snippet so it records the number of output tokens a candidate model produced
# Example: Track token usage
response = model.generate(prompt)
input_tokens = response.usage.prompt_tokens
output_tokens = response.usage. ?
total_tokens = response.usage.total_tokensThe usage object reports input as prompt_tokens and output as completion_tokens. Tracking both lets you compare candidates on cost and against separate ITPM and OTPM limits.
Source: docs.databricks.comThe system-performance metrics: total token count (cost) and latency in seconds. Both are measured deterministically, so you can compare them directly. The candidate that meets your quality bar at lower cost and latency, within its rate limits, is the better fit.
Checkpoint 5 of 5· Exam question
A team has fine-tuned a foundation model and is preparing to move a customer-facing application to production, where consistent response latency and guaranteed throughput under load are required. How should the fine-tuned model be served?
Correct answer: A — On a provisioned throughput endpoint, since dedicated capacity delivers the consistent latency and throughput guarantees production workloads require
- On a provisioned throughput endpoint, since dedicated capacity delivers the consistent latency and throughput guarantees production workloads require. A provisioned throughput endpoint is correct because it reserves dedicated capacity for the fine-tuned model, which is what production workloads need to get predictable latency and throughput under load.
- On a pay-per-token endpoint, since shared multi-tenant capacity provides the same latency guarantees as dedicated endpoints at lower cost. This is incorrect because pay-per-token endpoints are multi-tenant and shared across customers, so first-token latency can degrade under contention and no throughput guarantee is provided.
- As the base, non-fine-tuned version of the model on a pay-per-token endpoint, since pay-per-token endpoints do not support custom fine-tuned weights. This correctly notes that pay-per-token endpoints only serve the curated foundation models rather than custom fine-tuned weights, but choosing to drop the fine-tuning entirely abandons the customization the team already invested in and still lacks throughput guarantees.
- Through a personal compute cluster notebook, since notebook-based serving scales automatically to match customer-facing production traffic. This is incorrect because a notebook attached to a personal compute cluster is meant for interactive development and does not provide the autoscaling, availability, or throughput guarantees needed for customer-facing production serving.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Priority pay-per-token is the right mode for any production workload, including ones with compliance requirements such as HIPAA.Why is that wrong?
Priority mode is for latency-sensitive traffic. Provisioned throughput is the mode recommended for production and the one listed with compliance certifications like HIPAA.
Covered in Match the serving mode to the workload
2.If a model appears in system.ai, it can be used from my workspace.Why is that wrong?
system.ai is a global listing. Whether you can query a model depends on workspace region, cross-geo settings and model availability.
Covered in Check that the model can carry the load
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“This mode is recommended for all production workloads, especially those that require high throughput, performance guarantees, fine-tuned models, or have additional security requirements.”
↩︎ Match the serving mode to the workload“Priority mode is opt-in per request, requires no commitment, and is billed at a higher per-token rate.”
↩︎ Match the serving mode to the workload“AI Functions optimized models: This mode is recommended for batch inference workloads.”
↩︎ Match the serving mode to the workload“Efficiently compare LLMs to see the best candidate for your use case, or swap a production model with a better performing one.”
↩︎ Confirm the choice by comparing candidates“Provisioned throughput endpoints are available with compliance certifications like HIPAA.”
↩︎ Exam trap 1“Provisioned throughput endpoints are available with compliance certifications like HIPAA.”
↩︎ Prediction - 2.
“External models are third-party models hosted outside of Databricks.”
↩︎ Match the serving mode to the workload - 3.
“Smaller models for high-volume tasks: Use models like GPT OSS 20B for tasks that require higher throughput.”
↩︎ Check that the model can carry the load“Requests must be smaller than 200K input tokens or 400KB.”
↩︎ Check that the model can carry the load“create a provisioned throughput endpoint if the pay-per-token limits don't meet your use case requirements.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/machine-learning/model-serving/foundation-model-overviewOfficial docs
“Availability depends on your workspace region, cross-geo settings, and model availability.”
↩︎ Check that the model can carry the load“Seeing a model there does not mean it is available in your workspace.”
↩︎ Exam trap 2 - 5.https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/supported-modelsOfficial docs
“Databricks recommends customers evaluate their models to ensure they're operating as intended.”
↩︎ Confirm the choice by comparing candidates - 6.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/evaluate-assess-performanceOfficial docs
“Overall latency and token consumption are examples of chain performance metrics.”
↩︎ Confirm the choice by comparing candidates