CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 1 · Lesson 2/56

    Chaining Model Tasks with AI Functions: A Customer Review Pipeline

    Select model tasks to accomplish a given business requirement

    8 min read
    1.79% of exam
    2 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Decide when a task-specific AI Function fits and when ai_query is needed
    • Break a business requirement into an ordered chain of sentiment, classification, extraction and generation steps
    • Write the SQL that applies each model task to a table of reviews

    1.Task-specific function or ai_query?

    After you have named the model task, there is one more choice: a task-specific function or the general-purpose one. Databricks offers both kinds. Task-specific AI Functions are "Purpose-built functions optimized for a specific task, such as document parsing, entity extraction, classification, and sentiment analysis." The other kind is "ai_query — The general-purpose function for task and model flexibility. Provide a prompt and choose any supported Foundation Model API."

    The rule follows from what each one gives you. If the task is in the catalog (sentiment, classification, extraction, summarization, translation, masking), start with the task-specific function. You don't write a prompt, you don't parse output, and Databricks maintains the technique behind it. Use ai_query when the requirement needs a custom prompt or a particular model. The catalog says so itself next to its generation function: "For custom prompts or a specific model, see Use ai_query."

    All of these run where your data lives. They "can be run from Databricks SQL, notebooks, Lakeflow pipelines, and Workflows", with one exception: they are not available on Classic SQL warehouses.

    Checkpoint 1 of 5· Check yourself

    A requirement says replies must be generated by one specific model your security team has approved. Which function should you build on?

    Sources1

    2.One requirement, four chained tasks

    The Databricks customer-review example turns one vague requirement into a chain of tasks: deal with unhappy customers and improve the product descriptions. The pipeline determines the sentiment of each review. For negative reviews, it extracts information to classify the cause. It identifies whether the customer needs a response, and it generates a reply that mentions alternative products.

    Each step narrows the input for the next one. Sentiment is a cheap filter. Only negative reviews go on to cause classification, and the WHERE clause reuses the sentiment function to do the filtering:

    Classifying the cause of negative reviews against custom labels, after filtering by sentimentsql
    SELECT
      review,
      ai_classify(
        review,
        '["Arrives too late", "Wrong size", "Wrong color", "Dislike the style"]'
      ) AS reason
    FROM
      product_reviews
    WHERE
      ai_analyze_sentiment(review) = "negative"

    The labels come from the business. Logistics and merchandising care about different failure reasons, so they write the label list. That is what makes this classification rather than open-ended generation: the documentation notes that ai_classify() "is able to correctly categorize the negative reviews based on custom labels to allow for further analysis."

    To improve product descriptions, the team needs a specific fact from each review, such as the customer's usual size. That fact is a field, so the task is extraction. The same query also classifies fit, which shows that you can mix tasks in one SELECT.

    Checkpoint 2 of 5· Fill the gap

    This query pulls the customer's usual size out of each review and labels the fit. Which function completes it?

    SELECT
      review,
       ? (review, '["usual_size"]') AS usual_size,
      ai_classify(review, '["Size is wrong", "Size is right"]') AS fit
    FROM
      product_reviews

    Checkpoint 3 of 5· Put it in order

    Put the review pipeline's model tasks in the order the Databricks example applies them

    1. 1.Determine the sentiment of each review
    2. 2.Identify whether a response to the customer is required
    3. 3.For negative reviews, extract information to classify the cause
    4. 4.Generate a response mentioning alternative products

    Checkpoint 4 of 5· Exam question

    A media company receives thousands of incoming customer emails per day and needs each one automatically tagged with exactly one label from a fixed set: Billing, Technical Issue, Cancellation, or General Inquiry, so tickets route to the right team. Which model task best matches this requirement?

    Sources2

    3.Closing the loop: the generation task

    The last step is the only one that creates text the customer will read. That makes it generation, and the catalog's function for it is ai_gen, which will "Answer a user-provided prompt using a state-of-the-art AI model." Here, unlike the earlier steps, you write the prompt yourself. In the example, the prompt sets the length (60 words), the content (the customer's opinions are valued and a 30% discount coupon has been sent to their email), and the input (the review text joined onto the instruction). It runs only on rows where ai_analyze_sentiment(review) = "negative", the same filter the classification step used.

    The progression is worth remembering. Sentiment, classification and extraction each produce a narrow output. A label or a field can be checked, so those steps need no prompt. Generation produces open-ended text, so the business rules have to go into the prompt. If those rules ever require a model other than the one Databricks manages, the requirement has outgrown ai_gen. The catalog points you to ai_query for custom prompts or a specific model.

    Checkpoint 5 of 5· Check yourself

    Which step of the review pipeline is the one where you must write the instructions yourself in a prompt?

    Sources12

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Sentiment or label classification over text always needs ai_query with a carefully engineered prompt and output parsing.Why is that wrong?

      Task-specific functions such as ai_analyze_sentiment return the result directly. Databricks recommends starting with them because they need no customization.

      Covered in Task-specific function or ai_query?

    Practise it for real

    Turn the requirement "respond to unhappy customers and learn why they are unhappy" into chained AI Function calls over a product_reviews table

    1. 1.Run SELECT review, ai_analyze_sentiment(review) AS sentiment FROM product_reviews; from a notebook or a non-Classic SQL warehouse

      Why: Sentiment is the cheap first filter, and it decides which rows the later tasks process

      You should see: Each review is labeled positive, negative, neutral or mixed, with no prompt written

    2. 2.Run the ai_classify query with the labels Arrives too late, Wrong size, Wrong color and Dislike the style, filtered on negative sentiment

      Why: Cause classification uses labels the business defines, so the results can be grouped and counted

      You should see: A reason column holding one of your labels for each negative review

    3. 3.Run the ai_extract query for usual_size alongside ai_classify for fit

      Why: Extraction pulls out a named field that can drive changes to product descriptions

      You should see: usual_size and fit columns next to each review

    4. 4.Run ai_gen on negative reviews with a prompt for a 60-word reply that mentions the discount coupon

      Why: Generation is the only step whose output customers read, so its rules go into the prompt

      You should see: A reply column with a short, tailored response for each negative review

    Stuck? Get a nudge

    If a step's output needs a model or prompt the task-specific function can't give you, switch that step to ai_query and keep the rest of the chain unchanged.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Task-specific AI Functions — Purpose-built functions optimized for a specific task, such as document parsing, entity extraction, classification, and sentiment analysis.”
      ↩︎ Task-specific function or ai_query?
      “AI Functions are not available on Classic SQL warehouses.”
      ↩︎ Task-specific function or ai_query?
      “Generate content. For custom prompts or a specific model, see Use ai_query:”
      ↩︎ Closing the loop: the generation task
      “Databricks recommends these functions for getting started because they invoke a state-of-the-art research techniques maintained by Databricks and do not require any customization.”
      ↩︎ Exam trap 1
      “ai_query — The general-purpose function for task and model flexibility. Provide a prompt and choose any supported Foundation Model API.”
      ↩︎ Checkpoint
      “Answer a user-provided prompt using a state-of-the-art AI model.”
      ↩︎ Checkpoint
    2. 2.
      “ai_classify() is able to correctly categorize the negative reviews based on custom labels to allow for further analysis.”
      ↩︎ One requirement, four chained tasks
      “You can find key information from a blob of text using ai_extract().”
      ↩︎ One requirement, four chained tasks
      “Mention their opinions are valued and a 30% discount coupon code has been sent to their email.”
      ↩︎ Closing the loop: the generation task
      “the function returns the sentiment for each review without any prompt engineering or parsing results.”
      ↩︎ Prediction
      “For negative reviews, extracts information from the review to classify the cause.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 5 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.