What you will be able to do
- Map a task to the right model type (general purpose, embeddings, vision, reasoning) using Databricks' task-type guidance
- Read a model entry's metadata (supported inputs, context window, parameters, reasoning behaviour) and eliminate models that cannot do the task
- Use operational metadata such as AI Playground support, retirement dates and model terms to avoid a choice that breaks later
- Explain why a metadata-based shortlist must still be evaluated before production
1.Start from the task type, not the model name
Model selection should start from the job, not from a famous model name. Databricks puts it simply: the model you choose directly affects the responses you get. For conversational responses you pick a chat model. For text embeddings you pick an embedding model. Databricks' foundation model documentation groups models into task types, and each type has recommended use cases. Read that grouping first. It removes most of the catalog before you open a single model card.
| Task type | Recommended use cases | Example Databricks-hosted endpoints |
|---|---|---|
| General purpose | Virtual assistants, customer support bots, interactive tutoring systems | databricks-gpt-oss-120b, databricks-meta-llama-3-3-70b-instruct, databricks-claude-sonnet-4-5 |
| Embeddings | Semantic search, retrieval augmented generation (RAG), topic clustering, sentiment analysis and text analytics | databricks-qwen3-embedding-0-6b, databricks-gte-large-en, databricks-bge-large-en |
| Vision | Object detection and recognition, image classification, image segmentation, document understanding | databricks-gemma-3-12b, databricks-llama-4-maverick, databricks-glm-5-3-flash |
| Reasoning | Code generation, content creation and summarization, agent orchestration | databricks-gpt-oss-20b, databricks-glm-5-3, databricks-deepseek-v4-flash-0731 |
Notice that one model can appear in several rows. For example, databricks-gemma-3-12b is listed under both general purpose and vision. The task type tells you which group of models to consider. It does not choose one for you. Also notice that the embeddings row is short. The only Databricks-hosted embedding endpoints listed are databricks-qwen3-embedding-0-6b, databricks-gte-large-en and databricks-bge-large-en. A chat model, however capable, is the wrong kind of model for producing embeddings.
Checkpoint 1 of 6· Match them up
Match each task type to the application it is recommended for.
Tap a term, then the definition that fits it.
Each task type's recommended use cases come from Databricks' foundation model types table. Embeddings are what semantic search and RAG retrieval rely on.
“Recommended for applications where semantic understanding, similarity comparison, and efficient retrieval or clustering of complex data are essential”Source: docs.databricks.com
2.Reading a model entry: inputs, context, size and reasoning
Each supported model has an entry on Databricks' detailed model list, and every entry uses the same pattern. You get an endpoint name, the supported inputs, and a description that typically gives context length, parameter count, intended workloads and reasoning behaviour. Some entries point to a fuller model card. For example, the DeepSeek V4.1 Flash entry sends you to its model card for architecture, evaluations and limitations. The entry gives you the short summary, and the card gives you the detail behind it.
Read the entries as a list of hard constraints. Supported inputs decides whether the model can see images at all. Context length decides whether your longest prompt fits. Parameter counts and "fast, cost-efficient" wording show where a model sits on the size-versus-cost spectrum. For OpenAI, Google Gemini and Anthropic models, the context window and maximum output tokens match the values the provider publishes.
| Endpoint name | Supported inputs | Context / size metadata | Other metadata to note |
|---|---|---|---|
| databricks-glm-5-3 | text | Up to 1,048,576 tokens context; up to 65,536 output tokens | Reasoning is always enabled; supports function calling and structured output |
| databricks-glm-5-3-flash | text, image | 1,048,576 tokens; 320 billion total, 18 billion active parameters | Always reasons; reasoning cannot be disabled |
| databricks-gpt-oss-20b | text | 128K token context window | Lightweight; excels at real-time copilots and batch inference |
| databricks-gemma-3-12b | text, image | Up to 128K token context; 12 billion parameters | Multilingual support for over 140 languages; Gemma 3 terms apply |
| databricks-deepseek-v4-flash-0731 | text | 284 billion total, 13 billion active parameters | Optimized for fast, cost-efficient reasoning, coding and agentic tool use |
Checkpoint 2 of 6· Check yourself
A retailer needs a model that answers customer questions about uploaded product photos in many languages. Based only on the metadata above, which endpoint fits?
Gemma 3 12B accepts text and image inputs and lists multilingual support for over 140 languages. The other three list text as their only supported input.
“Gemma 3 has up to a 128K token context and provides multilingual support for over 140 languages.”Source: docs.databricks.com
Checkpoint 3 of 6· Exam question
An ML team plans to fine-tune a pretrained foundation model from the `system.ai` Unity Catalog schema on their internal support-ticket transcripts, then deploy the resulting model as a customer-facing chat assistant. Per the model card guidance for these pre-installed models, which variant should they select as the starting point for fine-tuning?
Correct answer: C — Base version
- A. The instruct variant has already been aligned with instruction-following and conversational behavior, so it is the recommended starting point for direct deployment rather than for further fine-tuning on custom data.
- B. A quantized variant trades numerical precision for lower memory footprint and faster inference; model cards do not recommend it as the preferred starting point for a fine-tuning workflow, since reduced precision can degrade training stability.
- C. The base variant is correct because Databricks model card guidance recommends starting from the base version when the plan is to further train the model on domain-specific data such as support-ticket transcripts, reserving the instruct version for direct deployment.
- D. A vision-language variant is built to accept image and text inputs together; the support-ticket transcripts described here are text-only, so this variant does not match the task and is not the guided starting point.
Sources3
3.Operational metadata: tooling, retirement and terms
Not every line in a model entry describes capability. Some lines describe how and how long you can use the model, and missing them causes problems later.
Tooling support. Several entries, such as GPT-6 Sol and GPT-5.5, say the model is not supported in AI Playground and must be called through the Responses API. If your team plans to test prompts in the Playground, that limits which models you can try there.
Lifecycle. Entries carry retirement notices. DeepSeek V4 Pro (0813) will be retired on October 30, 2026, and the entry links to a replacement model and migration guidance. A model that is about to retire is a poor choice for a new application, however well it fits the task.
Availability and terms. Entries state how a model is offered, for example "Pay-per-token" with a link to supported regions. They also say customers are responsible for complying with applicable model terms. Some families carry their own terms. For example, Gemma 3 has its own terms and Acceptable Use Policy.
Checkpoint 4 of 6· Check yourself
It is October 2026. Two text-only DeepSeek models both meet your accuracy needs for a new agentic coding assistant. One entry carries a retirement notice for October 30, 2026. What should you do?
The retirement notice is lifecycle metadata. Building a new application on a model that retires this month guarantees an immediate migration, and the entry itself points to a recommended replacement.
“DeepSeek V4 Pro (0813) will be retired on October 30, 2026.”Source: docs.databricks.com
Checkpoint 5 of 6· Exam question
A startup is building a commercial product on top of an open model listed in Databricks Marketplace and wants to avoid legal exposure from redistribution or usage restrictions. Before selecting the model, which model card field should they review first?
Correct answer: C — License and use policy
- A. Benchmark evaluation scores such as MMLU or HELM indicate the model's quality on standard tasks, but they carry no information about whether commercial redistribution or usage is legally permitted.
- B. How many times a model has been downloaded from the Marketplace reflects popularity, not the legal terms under which the model may be used, so it does not mitigate the startup's licensing risk.
- C. The license and acceptable use policy is correct because it specifies exactly what commercial, redistribution, and modification rights the model grants, which is the direct control point for avoiding legal exposure before building a product on top of the model.
- D. The supported context window describes how much text the model can process per request and is a functional characteristic, not a legal one, so checking it does not address redistribution or usage risk.
Sources3
4.Metadata produces a shortlist, not a final answer
Every model entry on the supported-models page ends with the same warning. AI models can make mistakes, misunderstand intent and hallucinate, and Databricks recommends that customers evaluate their models to make sure they work as intended. Metadata tells you which models *can* do the task: the right task type, the right inputs, enough context, available in your region, not retiring, and with terms you can meet. It cannot tell you which of those models does the task best on your data. Databricks presents the Foundation Model APIs as a way to compare LLMs efficiently and find the best candidate for your use case. Run that comparison on the shortlist that the metadata produced.
Checkpoint 6 of 6· Check yourself
Your metadata screen leaves two models that both fit the task type, inputs, context length, region and terms. What does Databricks recommend before you rely on either one?
Metadata only shows which models can do the task. Databricks recommends that customers evaluate their models to ensure they operate as intended, and this applies to every model.
“Databricks recommends customers evaluate their models to ensure they're operating as intended.”Source: docs.databricks.com
Because the warning applies to every model. Any model can hallucinate or misunderstand intent. A strong model card does not exempt a model from evaluation on your own use case.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Any model in the supported-models list can be tried in AI Playground before you pick it.Why is that wrong?
Some entries say explicitly that the model is not supported in AI Playground and must be called through the Responses API.
Covered in Operational metadata: tooling, retirement and terms
2.Two models from the same family with the same context window accept the same inputs.Why is that wrong?
Supported inputs are listed per model. GLM-5.3 is text-only, while GLM-5.3-Flash accepts text and image inputs, even though both list a 1,048,576-token context.
Covered in Reading a model entry: inputs, context, size and reasoning
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“The model you choose directly affects the results of the responses you get from the API calls.”
↩︎ Start from the task type, not the model name“for generating embeddings of text, you can choose an embedding model.”
↩︎ Start from the task type, not the model name - 2.https://docs.databricks.com/aws/en/machine-learning/model-serving/score-foundation-modelsOfficial docs
“Recommended for scenarios where natural, multi-turn dialogue and contextual understanding are needed”
↩︎ Start from the task type, not the model name“Recommended for applications where semantic understanding, similarity comparison, and efficient retrieval or clustering of complex data are essential”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/machine-learning/foundation-model-apis/supported-modelsOfficial docs
“For model architecture, evaluations, and limitations, see the DeepSeek V4.1 Flash model card.”
↩︎ Reading a model entry: inputs, context, size and reasoning“It supports function calling, parallel tool calls, structured output, a context window of up to 1,048,576 tokens, and up to 65,536 output tokens.”
↩︎ Reading a model entry: inputs, context, size and reasoning“For OpenAI, Google Gemini, and Anthropic models, the context window and maximum output tokens match the values published by the respective model provider.”
↩︎ Reading a model entry: inputs, context, size and reasoning“Customers are responsible for ensuring their compliance with applicable model terms.”
↩︎ Operational metadata: tooling, retirement and terms“This model is not supported in AI Playground. Use the Responses API to interact with this model.”
↩︎ Operational metadata: tooling, retirement and terms“Databricks recommends customers evaluate their models to ensure they're operating as intended.”
↩︎ Metadata produces a shortlist, not a final answer“This model is not supported in AI Playground. Use the Responses API to interact with this model.”
↩︎ Exam trap 1“GLM-5.3 is a text-only mixture of experts (MoE) language model developed by Zhipu AI for coding and agentic tool use.”
↩︎ Exam trap 2“GLM-5.3 is a text-only mixture of experts (MoE) language model developed by Zhipu AI for coding and agentic tool use.”
↩︎ Prediction“Gemma 3 has up to a 128K token context and provides multilingual support for over 140 languages.”
↩︎ Checkpoint“DeepSeek V4 Pro (0813) will be retired on October 30, 2026.”
↩︎ Checkpoint - 4.
“Efficiently compare LLMs to see the best candidate for your use case”
↩︎ Metadata produces a shortlist, not a final answer