What you will be able to do
- Choose between a chat model and an embedding model based on the task the application performs
- Use Databricks model descriptions to place a model on the frontier, balanced or cost-efficient tier of its family
- Rule out models that break a hard requirement: input modality, sampling parameters, data retention or tooling
Key concept
Use-case fit — No model is best for every application. You pick the LLM whose task type, input modalities, capability tier and operating constraints fit what your application actually needs, because that choice determines the responses you get.
1.Start from the application, not from the model list
Databricks gives you a long list of hosted models. The Unity Gateway list covers Anthropic, OpenAI, Google, Meta, Alibaba Cloud, xAI, Moonshot AI, Zhipu AI, DeepSeek and others, and through external models you can also reach third-party providers. With that many options, a better question than "which model is strongest?" is "which model fits what this application needs?" The external models documentation says it directly: "The model you choose directly affects the results of the responses you get from the API calls."
The first attribute to settle is the task type. An application that holds a conversation, answers questions or calls tools needs a chat (generative) model. An application that turns text into vectors for search needs an embedding model. These are separate kinds of model, and the endpoint types reflect that: Model Serving distinguishes llm/v1/completions, llm/v1/chat and llm/v1/embeddings. Among the hosted models, entries such as Qwen3-Embedding-0.6B and GTE Large (En) are embedding models, listed separately from the large language models in the rate-limit tables.
Once you know the task type, each further attribute (inputs, depth of reasoning, cost sensitivity, data handling) narrows the list. Some attributes only make a model a weaker choice. Others rule it out entirely, and those are covered in the last section of this page.
One gap to be aware of: these sources name the embedding endpoints but do not describe how to choose an embedding model by its context length. Don't treat this page as covering that.
Checkpoint 1 of 4· Check yourself
A team needs to vectorize product manuals for a similarity-search index. According to the external models guidance, which kind of model should they pick?
The task decides the model type. Generating embeddings calls for an embedding model, and conversational responses call for a chat model.
“for generating conversational responses, you can choose a chat model. Conversely, for generating embeddings of text, you can choose an embedding model.”Source: docs.databricks.com
Sources1
2.Read the model descriptions as a tier map
Within a single provider family, the supported-models page describes each model by what it is for. Read side by side, the descriptions sort models into tiers: a frontier model for the hardest work, a balanced model for everyday work, and a fast, cost-efficient model for volume. For example, the docs call GPT-6 Astra OpenAI's "frontier model for enterprise reasoning, structured document processing, coding, agentic search, multimodal inputs, and long-context workloads." GPT-6 Sol is described as strong on agentic coding and complex tool use "at a lower cost than GPT-6 Astra." GPT-6 Luna sits at the cheapest end of the family.
| Endpoint name | Supported inputs | Positioning in the docs |
|---|---|---|
| databricks-gpt-6-astra | text, image | Frontier model: enterprise reasoning, document processing, long-context workloads |
| databricks-gpt-6-luna | text, image | Fast and cost-efficient; lowest cost in the GPT-6 lineup |
| databricks-gpt-5-6-terra | text, image | Balanced; everyday agentic and reasoning work at lower cost than GPT-5.5 |
| databricks-gpt-5-5-pro | text, image | Higher-accuracy variant for deep research, advanced math, high-stakes reasoning |
| databricks-claude-sonnet-5 | text, image | Near-Opus-level intelligence with Sonnet cost efficiency and speed |
| databricks-claude-fable-5 | text | Long-running, autonomous knowledge work and coding |
Reasoning depth is a second dial, separate from the tier. Hybrid models such as Claude Sonnet 4.6 offer "two modes: near-instant responses and extended thinking for deeper reasoning based on the complexity of the task." That fits applications where some requests are simple and others are hard. Other models always reason and expose effort levels instead. Claude Opus 5.5 accepts low, medium, high, xhigh and max effort levels; the docs say "Databricks uses medium by default, and reasoning cannot be disabled." Claude Fable 5.1 likewise uses always-on adaptive thinking with per-turn effort controls. When an application needs fast responses for simple requests, look for a model that has a near-instant mode, or run an always-reasoning model at low effort.
Checkpoint 2 of 4· Match them up
Match each application need to the model whose documented positioning fits it best
Tap a term, then the definition that fits it.
Each model description states its intended workload. Pro targets the hardest problems, Luna targets lowest cost, Terra is the balanced option, and Sonnet 4.6 is a hybrid with two response modes.
“GPT-5.5 Pro is a higher-accuracy variant of GPT-5.5 aimed at the hardest problems, including deep research, advanced math, and high-stakes reasoning.”Source: docs.databricks.com
Checkpoint 3 of 4· Exam question
A company is building a high-volume customer support chatbot that only needs to answer simple, repetitive FAQ-style questions drawn from a fixed knowledge base. Response latency and per-request cost at scale are the team's top concerns. Which approach best matches the LLM to this application's attributes?
Correct answer: A — Deploy a small foundation model with low latency and low cost per token, since the task requires simple pattern matching across FAQ content rather than complex reasoning
- Deploy a small foundation model with low latency and low cost per token, since the task requires simple pattern matching across FAQ content rather than complex reasoning. A small, low-latency model is correct because the task complexity here is low and does not require the deeper reasoning capacity of a large model, so matching model size to task difficulty minimizes both latency and cost at scale.
- Deploy the largest available foundation model with the longest context window, since more parameters always produce more accurate answers regardless of task complexity. This is incorrect because parameter count and context length are only two of several relevant attributes, and using the largest model for a simple, high-volume task increases cost and latency without a meaningful quality benefit for this use case.
- Deploy a mid-size model fine-tuned for code generation, since fine-tuning on any task consistently reduces inference latency for conversational workloads. This is incorrect because fine-tuning for an unrelated task like code generation does not transfer to conversational FAQ answering, and fine-tuning itself does not inherently reduce serving latency.
- Deploy a model optimized for multilingual translation, since translation-focused models generalize best to customer support conversations in a single language. This is incorrect because a translation-specialized model is optimized for a different capability and offers no particular advantage for single-language FAQ conversations.
Sources2
3.Constraints that rule a model out
Some attributes are pass/fail, not trade-offs. Check these before comparing quality.
Input modality. Each endpoint lists its supported inputs. Most of the models above accept text, image, but Claude Fable 5 lists only text. If your application sends images, such as screenshots or scanned forms, a text-only model won't work no matter how capable it is.
Request parameters. If an application depends on a sampling parameter, the model has to accept it. The docs say: "Claude Sonnet 5 does not support the temperature, top_p, or top_k sampling parameters. Requests that include these parameters return a 400 error." Defaults also matter: the limits page advises always setting max_tokens for Claude Sonnet 4 "to avoid the default 1,000 token limit."
Data handling. For Claude Fable 5.1 and Claude Fable 5, prompts and responses are retained for 30 days for trust and safety purposes. The docs are explicit: "Customers who opt out of data retention cannot use Claude Fable 5.1." If an application's data policy requires opting out of retention, these models are excluded.
Tooling and context window. Several OpenAI models note "This model is not supported in AI Playground. Use the Responses API to interact with this model." So if you plan to prototype in the Playground, check this first. For context length, the page says that for OpenAI, Google Gemini and Anthropic models, "the context window and maximum output tokens match the values published by the respective model provider", so you look those limits up with the provider.
Checkpoint 4 of 4· Check yourself
A healthcare startup's policy requires opting out of all provider data retention. They want the most autonomous agent model on the list and lean toward Claude Fable 5.1. What do the docs say?
Fable 5.1 keeps prompts and responses for 30 days for trust and safety, and opting out of retention makes the model unavailable to you. That is a hard constraint, not a trade-off.
“Customers who opt out of data retention cannot use Claude Fable 5.1.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.The most powerful frontier model is always the right pick, because higher capability means better results for any application.Why is that wrong?
Selection is about fit. The docs position cheaper models for strong everyday work, and they recommend Opus 5 for cost-sensitive work precisely because it keeps near-top quality at low effort.
Covered in Read the model descriptions as a tier map
2.Any hosted chat model accepts the standard sampling parameters, so swapping models never breaks request code.Why is that wrong?
Claude Sonnet 5 rejects temperature, top_p and top_k with a 400 error, so parameter support is a hard constraint on which model you can choose.
Covered in Constraints that rule a model out
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The model you choose directly affects the results of the responses you get from the API calls.”
↩︎ Start from the application, not from the model list“for generating conversational responses, you can choose a chat model. Conversely, for generating embeddings of text, you can choose an embedding model.”
↩︎ Start from the application, not from the model list“Therefore, choose a model that fits your use-case requirements.”
↩︎ Key concept - 2.https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/supported-modelsOfficial docs
“GPT-6 Astra is OpenAI's frontier model for enterprise reasoning, structured document processing, coding, agentic search, multimodal inputs, and long-context workloads.”
↩︎ Read the model descriptions as a tier map“GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 family, bringing strong reasoning and agentic capability at the lowest cost in the lineup.”
↩︎ Read the model descriptions as a tier map“It offers two modes: near-instant responses and extended thinking for deeper reasoning based on the complexity of the task.”
↩︎ Read the model descriptions as a tier map“Databricks uses medium by default, and reasoning cannot be disabled.”
↩︎ Read the model descriptions as a tier map“Claude Sonnet 5 does not support the temperature, top_p, or top_k sampling parameters.”
↩︎ Constraints that rule a model out“the context window and maximum output tokens match the values published by the respective model provider.”
↩︎ Constraints that rule a model out“it delivers near-top quality at low and medium reasoning effort, making it well suited for long-context, cost-sensitive enterprise workflows.”
↩︎ Exam trap 1“Requests that include these parameters return a 400 error.”
↩︎ Exam trap 2“it delivers near-top quality at low and medium reasoning effort, making it well suited for long-context, cost-sensitive enterprise workflows.”
↩︎ Prediction“GPT-5.5 Pro is a higher-accuracy variant of GPT-5.5 aimed at the hardest problems, including deep research, advanced math, and high-stakes reasoning.”
↩︎ Checkpoint“Customers who opt out of data retention cannot use Claude Fable 5.1.”
↩︎ Checkpoint - 3.
“Always specify max_tokens when using Claude Sonnet 4 to avoid the default 1,000 token limit”
↩︎ Constraints that rule a model out