CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 1 · Lesson 2/56

    Choosing a Model Task for a Business Requirement on Databricks

    Select model tasks to accomplish a given business requirement

    10 min read
    1.79% of exam
    7 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Turn a business requirement into a named model task: summarization, classification, extraction, generation, or chat
    • Map each task to the matching task-specific AI Function
    • Choose between the general-purpose, embeddings, vision and reasoning foundation model types
    • Tell when a requirement fits Knowledge Assistant and when it fits Information Extraction

    Key concept

    Model task — A model task is the kind of transformation a requirement asks for, such as classifying text, extracting fields, summarizing, generating new content, or holding a conversation. On Databricks you pick the task first, because it decides which purpose-built function, model type or Agent Bricks product fits. Picking a model comes after that.

    1.Start from the output the business needs

    Business requirements rarely name a model task. A stakeholder describes an outcome, and you have to work out what kind of transformation produces it. The Databricks ML lifecycle guidance puts the same question first for any model: "What is the prediction target, and what class of ML problem does that imply: classification, regression, forecasting, recommendation, ranking, anomaly detection, or something else?" Generative AI requirements need the same step. You ask what the output looks like.

    - A label chosen from a fixed list (complaint reason, sentiment, ticket category) is classification. - Named fields pulled out of text (contract parties, invoice line items, a customer's usual size) is extraction. - A shorter version of the same content is summarization. - New text that did not exist before (a reply to a customer, a product description) is generation. - A multi-turn exchange where the user asks follow-up questions is chat.

    The unhappy-customers requirement from the prediction above is really several tasks. First you classify sentiment, then you classify the cause, and maybe you also generate a reply. Splitting a requirement into its tasks is the skill this objective tests.

    Checkpoint 1 of 6· Check yourself

    After flagging negative reviews, a team wants a personalized reply drafted for each unhappy customer. Which model task is that?

    Sources12

    2.From task to task-specific AI Function

    Once you know the task, Databricks often has a built-in function for it. AI Functions are grouped by task, and the documentation says why they are the default starting point: "Databricks recommends these functions for getting started because they invoke a state-of-the-art research techniques maintained by Databricks and do not require any customization." The table maps common requirements to these functions. Look at the inputs each one needs. ai_classify takes the labels you provide at query time, and ai_extract takes a schema you define.

    Business requirement → model task → task-specific AI Function
    Business requirementModel taskAI FunctionWhat the docs say it does
    Tag each review as positive, negative, neutral or mixedClassification (sentiment)ai_analyze_sentimentPerform sentiment analysis on input text
    Route complaints into reasons you defineClassificationai_classifyClassify input text according to labels you provide
    Pull named fields out of documents or textExtractionai_extractExtract structured fields from documents or text using a schema you define
    Condense long textSummarizationai_summarizeGenerate a summary of text
    Hide personal details before sharing textTransformation (masking)ai_maskMask specified entities in text
    Draft new text from an instructionGenerationai_genAnswer a user-provided prompt
    Turn PDFs into text, tables and layoutDocument parsingai_parse_documentParse structured content (text, tables, figure descriptions) and layout from unstructured documents

    Checkpoint 2 of 6· Check yourself

    A finance team wants every invoice PDF turned into rows with vendor, due date and payment terms, using column names they choose. Which model task and function fit best?

    Checkpoint 3 of 6· Exam question

    A legal operations team wants to pull specific fields (contract start date, renewal term, counterparty name) out of thousands of scanned PDF contracts and load the results into a Delta table with a fixed schema. Which model task should the application be built around?

    Sources3

    3.Foundation model task types: chat, embeddings, vision, reasoning

    When you call foundation models directly instead of through a task-specific function, the task still guides your choice. Databricks groups its foundation models into four task types, and each has its own recommended use cases. The most common mistake is mixing up the first two. A general-purpose model *talks*. An embedding model returns a vector, not an answer: it transforms "complex data—such as text, images, or audio—into compact numerical vectors called embeddings." That makes it the right choice for retrieval and similarity, and the wrong choice for anything a user reads directly.

    Foundation model task types and the requirements they fit
    Task typeWhat it producesRecommended use casesExample Databricks-hosted model
    General purposeContextual responses in natural, multi-turn conversationVirtual assistants, customer support bots, interactive tutoring systemsdatabricks-claude-sonnet-4-5
    EmbeddingsCompact numerical vectorsSemantic search, RAG, topic clustering, sentiment analysis and text analyticsdatabricks-gte-large-en
    VisionInterpretation of images and videoObject detection and recognition, image classification, image segmentation, document understandingdatabricks-gemma-3-12b
    ReasoningStep-by-step analysis that breaks down tasksCode generation, content creation and summarization, agent orchestrationdatabricks-gpt-oss-120b

    Checkpoint 4 of 6· Match them up

    Match each business requirement to the foundation model task type Databricks recommends for it

    Tap a term, then the definition that fits it.

    Checkpoint 5 of 6· Exam question

    A support organization has a large internal knowledge base of product manuals and wants employees to ask natural-language questions and receive grounded answers with citations back to the source manuals. Which Databricks Agent Bricks offering is the most direct fit for this business requirement?

    Sources4

    4.Two tasks as products: Knowledge Assistant and Information Extraction

    Agent Bricks packages two of these tasks as products. They answer different requirements, so match them to the output you need.

    Knowledge Assistant handles *chat over your documents*. Its purpose is to "create a chatbot that can answer questions about your documents and provide high-quality responses with citations." The documented use cases include answering questions from product documentation, employee HR questions and customer inquiries from support knowledge bases. It also takes prompt engineering off your hands: "Knowledge Assistant automates prompt engineering under the hood, based on your data and feedback."

    Information Extraction handles *structured output from unstructured input*. It "transforms unstructured documents and text into key, structured insights using a defined schema." Examples include pulling legal parties and terms out of contracts, line items and payment terms out of invoices, and key details out of medical records. It is "available through the Agent Bricks UI, in SQL, and with the REST API", and it relies on the ai_extract SQL function.

    Here is a quick way to tell them apart. If a person will read the answer to a question, the task is chat, so use Knowledge Assistant. If a downstream table or application consumes named fields, the task is extraction, so use Information Extraction. The sources describe what each product is for. They do not give a general rule for when to build a custom agent instead.

    Checkpoint 6 of 6· Check yourself

    HR wants employees to ask free-form questions about leave policy and get answers that cite the policy documents. Which Agent Bricks product fits this requirement?

    Sources567

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.ai_classify trains a classifier on your historical data before it can label text.Why is that wrong?

      ai_classify sorts text into labels you provide when you call it, with no training. ai_predict_class (Beta) is the function that trains a model and fills in NULL target values.

      Covered in From task to task-specific AI Function

    2. 2.An embedding model can serve as the conversational model behind a customer support bot.Why is that wrong?

      Embedding models output numerical vectors for search, retrieval and clustering. Multi-turn support bots are the recommended use of the general-purpose task type.

      Covered in Foundation model task types: chat, embeddings, vision, reasoning

    3. 3.Knowledge Assistant and Information Extraction are interchangeable ways to work with documents.Why is that wrong?

      Information Extraction produces structured fields from a defined schema for downstream use. Knowledge Assistant answers questions in a chat, with citations.

      Covered in Two tasks as products: Knowledge Assistant and Information Extraction

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “What is the prediction target, and what class of ML problem does that imply: classification, regression, forecasting, recommendation, ranking, anomaly detection, or something else?”
      ↩︎ Start from the output the business needs
    2. 2.
      “you can use the ai_gen() function to generate a response to a customer based on their complaint”
      ↩︎ Start from the output the business needs
    3. 3.
      “Databricks recommends these functions for getting started because they invoke a state-of-the-art research techniques maintained by Databricks and do not require any customization.”
      ↩︎ From task to task-specific AI Function
      “Classify input text according to labels you provide using state-of-the-art research techniques.”
      ↩︎ From task to task-specific AI Function
      “Task-specific AI Functions — Purpose-built functions optimized for a specific task, such as document parsing, entity extraction, classification, and sentiment analysis.”
      ↩︎ Key concept
      “Train a classification model and predict class labels for rows where the target column is NULL.”
      ↩︎ Exam trap 1
      “Extract structured fields from documents or text using a schema you define.”
      ↩︎ Checkpoint
    4. 4.
      “Object detection and recognition Image classification Image segmentation Document understanding”
      ↩︎ Foundation model task types: chat, embeddings, vision, reasoning
      “Code generation Content creation and summarization Agent orchestration”
      ↩︎ Foundation model task types: chat, embeddings, vision, reasoning
      “Embedding models are machine learning systems that transform complex data—such as text, images, or audio—into compact numerical vectors called embeddings.”
      ↩︎ Exam trap 2
      “Semantic search Retrieval augmented generation (RAG) Topic clustering Sentiment analysis and text analytics”
      ↩︎ Prediction
      “Recommended for scenarios where natural, multi-turn dialogue and contextual understanding are needed: Virtual assistants Customer support bots Interactive tutoring systems.”
      ↩︎ Checkpoint
    5. 5.
      “Answer employee questions related to HR policies.”
      ↩︎ Two tasks as products: Knowledge Assistant and Information Extraction
      “Use Knowledge Assistant to create a chatbot that can answer questions about your documents and provide high-quality responses with citations.”
      ↩︎ Checkpoint
    6. 6.
      “Knowledge Assistant automates prompt engineering under the hood, based on your data and feedback.”
      ↩︎ Two tasks as products: Knowledge Assistant and Information Extraction
    7. 7.
      “Information extraction is available through the Agent Bricks UI, in SQL, and with the REST API.”
      ↩︎ Two tasks as products: Knowledge Assistant and Information Extraction
      “Information Extraction transforms unstructured documents and text into key, structured insights using a defined schema.”
      ↩︎ Exam trap 3

    Continue to page 2 of 2

    Chaining Model Tasks with AI Functions: A Customer Review Pipeline

    Spotted a mistake, or was something unclear? Tell us.