What you will be able to do
- Turn a business requirement into a named model task: summarization, classification, extraction, generation, or chat
- Map each task to the matching task-specific AI Function
- Choose between the general-purpose, embeddings, vision and reasoning foundation model types
- Tell when a requirement fits Knowledge Assistant and when it fits Information Extraction
Key concept
Model task — A model task is the kind of transformation a requirement asks for, such as classifying text, extracting fields, summarizing, generating new content, or holding a conversation. On Databricks you pick the task first, because it decides which purpose-built function, model type or Agent Bricks product fits. Picking a model comes after that.
1.Start from the output the business needs
Business requirements rarely name a model task. A stakeholder describes an outcome, and you have to work out what kind of transformation produces it. The Databricks ML lifecycle guidance puts the same question first for any model: "What is the prediction target, and what class of ML problem does that imply: classification, regression, forecasting, recommendation, ranking, anomaly detection, or something else?" Generative AI requirements need the same step. You ask what the output looks like.
- A label chosen from a fixed list (complaint reason, sentiment, ticket category) is classification. - Named fields pulled out of text (contract parties, invoice line items, a customer's usual size) is extraction. - A shorter version of the same content is summarization. - New text that did not exist before (a reply to a customer, a product description) is generation. - A multi-turn exchange where the user asks follow-up questions is chat.
The unhappy-customers requirement from the prediction above is really several tasks. First you classify sentiment, then you classify the cause, and maybe you also generate a reply. Splitting a requirement into its tasks is the skill this objective tests.
Checkpoint 1 of 6· Check yourself
After flagging negative reviews, a team wants a personalized reply drafted for each unhappy customer. Which model task is that?
The reply is new text that did not exist before, so the task is generation. The Databricks review example uses ai_gen for exactly this step.
“you can use the ai_gen() function to generate a response to a customer based on their complaint”Source: docs.databricks.com
2.From task to task-specific AI Function
Once you know the task, Databricks often has a built-in function for it. AI Functions are grouped by task, and the documentation says why they are the default starting point: "Databricks recommends these functions for getting started because they invoke a state-of-the-art research techniques maintained by Databricks and do not require any customization." The table maps common requirements to these functions. Look at the inputs each one needs. ai_classify takes the labels you provide at query time, and ai_extract takes a schema you define.
| Business requirement | Model task | AI Function | What the docs say it does |
|---|---|---|---|
| Tag each review as positive, negative, neutral or mixed | Classification (sentiment) | ai_analyze_sentiment | Perform sentiment analysis on input text |
| Route complaints into reasons you define | Classification | ai_classify | Classify input text according to labels you provide |
| Pull named fields out of documents or text | Extraction | ai_extract | Extract structured fields from documents or text using a schema you define |
| Condense long text | Summarization | ai_summarize | Generate a summary of text |
| Hide personal details before sharing text | Transformation (masking) | ai_mask | Mask specified entities in text |
| Draft new text from an instruction | Generation | ai_gen | Answer a user-provided prompt |
| Turn PDFs into text, tables and layout | Document parsing | ai_parse_document | Parse structured content (text, tables, figure descriptions) and layout from unstructured documents |
Checkpoint 2 of 6· Check yourself
A finance team wants every invoice PDF turned into rows with vendor, due date and payment terms, using column names they choose. Which model task and function fit best?
The output is a set of named fields, which makes this extraction. ai_extract is built to pull those fields out using a schema you define.
“Extract structured fields from documents or text using a schema you define.”Source: docs.databricks.com
Checkpoint 3 of 6· Exam question
A legal operations team wants to pull specific fields (contract start date, renewal term, counterparty name) out of thousands of scanned PDF contracts and load the results into a Delta table with a fixed schema. Which model task should the application be built around?
Correct answer: A — Structured information extraction against a defined schema
- A. Pulling named fields such as dates, terms, and party names into a fixed schema is exactly what structured information extraction (e.g. `ai_extract`) is designed for, and it maps directly onto a Delta table with fixed columns.
- B. Generation would produce free-form prose summaries, which do not give the fixed, queryable columns (start date, term, counterparty) the team needs for the Delta table.
- C. A chat interface answers ad hoc questions interactively but does not batch-produce structured field values across thousands of documents for a table load.
- D. Classification assigns each document to one of a set of predefined labels, which does not recover the specific field values the team needs to populate table columns.
Sources3
3.Foundation model task types: chat, embeddings, vision, reasoning
When you call foundation models directly instead of through a task-specific function, the task still guides your choice. Databricks groups its foundation models into four task types, and each has its own recommended use cases. The most common mistake is mixing up the first two. A general-purpose model *talks*. An embedding model returns a vector, not an answer: it transforms "complex data—such as text, images, or audio—into compact numerical vectors called embeddings." That makes it the right choice for retrieval and similarity, and the wrong choice for anything a user reads directly.
| Task type | What it produces | Recommended use cases | Example Databricks-hosted model |
|---|---|---|---|
| General purpose | Contextual responses in natural, multi-turn conversation | Virtual assistants, customer support bots, interactive tutoring systems | databricks-claude-sonnet-4-5 |
| Embeddings | Compact numerical vectors | Semantic search, RAG, topic clustering, sentiment analysis and text analytics | databricks-gte-large-en |
| Vision | Interpretation of images and video | Object detection and recognition, image classification, image segmentation, document understanding | databricks-gemma-3-12b |
| Reasoning | Step-by-step analysis that breaks down tasks | Code generation, content creation and summarization, agent orchestration | databricks-gpt-oss-120b |
Checkpoint 4 of 6· Match them up
Match each business requirement to the foundation model task type Databricks recommends for it
Tap a term, then the definition that fits it.
Each pairing follows the recommended use cases listed for that task type: multi-turn dialogue for general purpose, clustering for embeddings, document understanding for vision, and agent orchestration for reasoning.
“Recommended for scenarios where natural, multi-turn dialogue and contextual understanding are needed: Virtual assistants Customer support bots Interactive tutoring systems.”Source: docs.databricks.com
Checkpoint 5 of 6· Exam question
A support organization has a large internal knowledge base of product manuals and wants employees to ask natural-language questions and receive grounded answers with citations back to the source manuals. Which Databricks Agent Bricks offering is the most direct fit for this business requirement?
Correct answer: A — Knowledge Assistant, since it is built for grounded document question-answering over a source corpus
- A. Knowledge Assistant is purpose-built for retrieving and answering questions grounded in a document corpus, including citations back to source material, which matches this single-domain Q&A requirement.
- B. Supervisor orchestrates multiple specialized agents and tools across different domains; a single knowledge base Q&A need does not require that multi-domain coordination layer.
- C. Information Extraction is designed to pull defined fields out of documents into structured records, not to answer open-ended natural-language questions with cited passages.
- D. Genie is designed for natural-language queries against structured tabular data, not for question-answering over unstructured product manuals.
Sources4
4.Two tasks as products: Knowledge Assistant and Information Extraction
Agent Bricks packages two of these tasks as products. They answer different requirements, so match them to the output you need.
Knowledge Assistant handles *chat over your documents*. Its purpose is to "create a chatbot that can answer questions about your documents and provide high-quality responses with citations." The documented use cases include answering questions from product documentation, employee HR questions and customer inquiries from support knowledge bases. It also takes prompt engineering off your hands: "Knowledge Assistant automates prompt engineering under the hood, based on your data and feedback."
Information Extraction handles *structured output from unstructured input*. It "transforms unstructured documents and text into key, structured insights using a defined schema." Examples include pulling legal parties and terms out of contracts, line items and payment terms out of invoices, and key details out of medical records. It is "available through the Agent Bricks UI, in SQL, and with the REST API", and it relies on the ai_extract SQL function.
Here is a quick way to tell them apart. If a person will read the answer to a question, the task is chat, so use Knowledge Assistant. If a downstream table or application consumes named fields, the task is extraction, so use Information Extraction. The sources describe what each product is for. They do not give a general rule for when to build a custom agent instead.
Checkpoint 6 of 6· Check yourself
HR wants employees to ask free-form questions about leave policy and get answers that cite the policy documents. Which Agent Bricks product fits this requirement?
The requirement is question answering with citations over documents, which is a chat task. Knowledge Assistant is built for this, and HR policy questions are one of its documented use cases.
“Use Knowledge Assistant to create a chatbot that can answer questions about your documents and provide high-quality responses with citations.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.ai_classify trains a classifier on your historical data before it can label text.Why is that wrong?
ai_classify sorts text into labels you provide when you call it, with no training. ai_predict_class (Beta) is the function that trains a model and fills in NULL target values.
Covered in From task to task-specific AI Function
2.An embedding model can serve as the conversational model behind a customer support bot.Why is that wrong?
Embedding models output numerical vectors for search, retrieval and clustering. Multi-turn support bots are the recommended use of the general-purpose task type.
Covered in Foundation model task types: chat, embeddings, vision, reasoning
3.Knowledge Assistant and Information Extraction are interchangeable ways to work with documents.Why is that wrong?
Information Extraction produces structured fields from a defined schema for downstream use. Knowledge Assistant answers questions in a chat, with citations.
Covered in Two tasks as products: Knowledge Assistant and Information Extraction
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“What is the prediction target, and what class of ML problem does that imply: classification, regression, forecasting, recommendation, ranking, anomaly detection, or something else?”
↩︎ Start from the output the business needs - 2.
“you can use the ai_gen() function to generate a response to a customer based on their complaint”
↩︎ Start from the output the business needs - 3.
“Databricks recommends these functions for getting started because they invoke a state-of-the-art research techniques maintained by Databricks and do not require any customization.”
↩︎ From task to task-specific AI Function“Classify input text according to labels you provide using state-of-the-art research techniques.”
↩︎ From task to task-specific AI Function“Task-specific AI Functions — Purpose-built functions optimized for a specific task, such as document parsing, entity extraction, classification, and sentiment analysis.”
↩︎ Key concept“Train a classification model and predict class labels for rows where the target column is NULL.”
↩︎ Exam trap 1“Extract structured fields from documents or text using a schema you define.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/machine-learning/model-serving/score-foundation-modelsOfficial docs
“Object detection and recognition Image classification Image segmentation Document understanding”
↩︎ Foundation model task types: chat, embeddings, vision, reasoning“Code generation Content creation and summarization Agent orchestration”
↩︎ Foundation model task types: chat, embeddings, vision, reasoning“Embedding models are machine learning systems that transform complex data—such as text, images, or audio—into compact numerical vectors called embeddings.”
↩︎ Exam trap 2“Semantic search Retrieval augmented generation (RAG) Topic clustering Sentiment analysis and text analytics”
↩︎ Prediction“Recommended for scenarios where natural, multi-turn dialogue and contextual understanding are needed: Virtual assistants Customer support bots Interactive tutoring systems.”
↩︎ Checkpoint - 5.
“Answer employee questions related to HR policies.”
↩︎ Two tasks as products: Knowledge Assistant and Information Extraction“Use Knowledge Assistant to create a chatbot that can answer questions about your documents and provide high-quality responses with citations.”
↩︎ Checkpoint - 6.https://docs.databricks.com/aws/en/agents/conceptsOfficial docs
“Knowledge Assistant automates prompt engineering under the hood, based on your data and feedback.”
↩︎ Two tasks as products: Knowledge Assistant and Information Extraction - 7.
“Information extraction is available through the Agent Bricks UI, in SQL, and with the REST API.”
↩︎ Two tasks as products: Knowledge Assistant and Information Extraction“Information Extraction transforms unstructured documents and text into key, structured insights using a defined schema.”
↩︎ Exam trap 3