CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 1 · Lesson 4/56

    Defining GenAI Pipeline Inputs and Outputs from Business Goals

    Translate business use case goals into a description of the desired inputs and outputs for the AI pipeline

    11 min read
    1.79% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Turn a business goal into concrete statements about pipeline input data, output surface, and serving requirements
    • Match a business requirement to a model task such as extraction, classification, parsing, or question answering
    • Specify a machine-readable output contract using response_format with json_schema or json_object, and know which schema features are unsupported

    Key concept

    Output contract — Before you choose any component, write down exactly what the pipeline takes in and what shape its answer must have. When code consumes the answer, the contract is usually a JSON structure enforced at generation time rather than free text.

    1.From a business goal to inputs, outputs, and constraints

    A business stakeholder rarely says "build me a retrieval chain." They say "cut the time support staff spend answering policy questions" or "get the invoice totals into our warehouse." Your first job is to restate that goal as a pipeline specification: what goes in, what comes out, and what limits the system has to meet. The Databricks ML lifecycle guidance frames this as scoping, done before anything is built, and its questions carry over to generative AI. Start with the target and the kind of problem it implies: "What is the prediction target, and what class of ML problem does that imply". For a GenAI pipeline, the answer is a task (summarise, classify, extract, generate, or converse), and that task decides most of what follows.

    Next, list the inputs that actually exist. Supporting data can be unstructured or structured. Unstructured inputs include PDFs, Office documents, wikis, images, and videos. Structured inputs include customer records in a warehouse, transaction data in a SQL database, and data from application APIs such as SAP or Salesforce. The difference matters because it decides how the pipeline will fetch context later: semantic search over chunks, or a SQL or API call. Then pin down the non-functional side: "What are the serving and production requirements: latency, throughput, and data freshness?" If the business needs answers to reflect documents that change every week, freshness is part of the input requirement.

    Checkpoint 1 of 4· Check yourself

    While scoping a pipeline, which set of items does the Databricks lifecycle guidance list as the serving and production requirements to pin down?

    Finally, decide what the output looks like and who consumes it. Databricks lists three forms a user-facing AI application can take: a chat app (for example, deployed with Databricks Apps), an API endpoint (for example, an agent deployed to Model Serving), and a SQL function for analysts (for example, an AI Function). A human reading a chat window can accept prose with citations. A downstream job writing to a Delta table needs fields with fixed names and types. These are two very different output contracts.

    Output surfaces named in the Databricks concepts guide and what each implies for the output design
    Output surfaceTypical consumerOutput design implication
    Chat app (Databricks Apps)A person in a conversationNatural-language answer; conversation history becomes part of the input
    API endpoint (agent on Model Serving)Another application or serviceResponse must be stable and parseable by the calling code
    SQL function (AI Function)An analyst running queries over tablesOne output value per row, ready to land in a column

    Sources123

    2.Matching the requirement to a model task

    Once the inputs and consumers are clear, name the task. Databricks recommends structured outputs for a group of business patterns that keep coming up: "Extracting data from large amounts of documents. For example, identifying and classifying product review feedback as negative, positive or neutral." Batch inference jobs whose outputs must follow a specified format, and general data processing that turns unstructured data into structured data, are on the same list. In each case the business goal is a table of fields or labels, not a paragraph.

    For document-heavy goals, Databricks Intelligent Document Processing (IDP) breaks the work into separate capabilities. Each one has a fixed input and output, and each is available in three forms: the Agent Bricks UI, a SQL AI Function, or a REST API. The Agent Bricks UI is the no-code place to try sample documents and refine a schema or label set. The SQL functions run the same capability at scale in Lakeflow pipelines. Matching the business phrasing to the right capability is often the whole design decision.

    IDP capabilities: business need mapped to input, output, and the SQL function that provides it
    Business needInputOutputSQL function
    Make raw files machine-readablePDFs, DOCX, images, PPTsStructured text, tables, figure descriptionsai_parse_document
    Pull named fields such as totals, dates, partiesDocuments or plain text plus a schema you defineStructured fieldsai_extract
    Route or tag documents by type or topicDocuments or text plus predefined categoriesA label (supports up to 500+ labels)ai_classify
    Feed a RAG or search applicationParsed documentsSemantic chunks with titles, section headers, page referencesai_prep_search (Beta)

    These stages chain together in a fixed direction. ai_parse_document produces a standardised bronze layer, and ai_extract and ai_classify then "operate directly on the parsed outputs". So the output of one stage is the declared input of the next. Write the pipeline description the same way: each stage gets a named input and a named output.

    Checkpoint 2 of 4· Check yourself

    A procurement team says: "From each supplier invoice PDF we need invoice number, due date, and total amount, using the field list we've agreed." Which capability matches the core requirement?

    Checkpoint 3 of 4· Exam question

    A retail company wants a pipeline that reads customer reviews and returns a single sentiment label (`positive`, `negative`, or `neutral`) together with a numeric confidence score, so a downstream dashboard can programmatically filter and chart results. Which combination of model task and output design best matches this business goal?

    Sources45

    3.Writing the output contract: response_format and JSON schemas

    When the consumer is code, describing the format in the prompt alone leaves room for drift. Databricks structured outputs let you put the contract into the request itself: "Specify your structured outputs using response_format in your chat request." Three output modes are available: "You can choose to generate text, unstructured JSON objects, and JSON objects that adhere to a specific JSON schema." Structured outputs are supported for chat models served through Foundation Model APIs, on both pay-per-token and provisioned throughput endpoints.

    The prompt and the schema do different jobs. In the research-paper extraction example, a system message gives the model its role and task, the user message carries the unstructured input, and a json_schema response_format named research_paper_extraction declares the fields (title, authors, abstract, keywords) with "strict": True.

    The system and user messages from the research-paper extraction example: the instruction prompt states the task, and the user message carries the raw inputpython
    messages = [{
            "role": "system",
            "content": "You are an expert at structured data extraction. You will be given unstructured text from a research paper and should convert it into the given structure."
          },
          {
            "role": "user",
            "content": "..."
          }]
    The chat completion call: the output contract travels with the request as response_formatpython
    response = client.chat.completions.create(
        model="databricks-gpt-oss-20b",
        messages=messages,
        response_format=response_format
    )

    Pick the mode based on how well you know the output shape. If downstream code expects specific fields, use json_schema. If you only need valid JSON and the fields cannot be fixed in advance, use json_object, which the docs illustrate as "an example of JSON extraction, but the JSON schema is not known before hand." In that example, the field names (name, size, price, color) appear only in the prompt text. Keep schemas simple, too. Foundation Model APIs support only a subset of JSON schema, because simpler schemas produce better JSON. The unsupported features are regex pattern, composition keywords (anyOf, oneOf, allOf, prefixItems, $ref), and lists of types other than [type, "null"].

    Checkpoint 4 of 4· Fill the gap

    This response_format asks for valid JSON when the field structure is not known in advance. Which type completes it?

    response_format = {
          "type": " ? ",
        }

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Setting response_format to json_object guarantees the output matches the fields my downstream table expects.Why is that wrong?

      json_object only guarantees valid JSON. It is meant for cases where the schema isn't known in advance. To enforce named, typed fields, use json_schema with a declared schema.

      Covered in Writing the output contract: response_format and JSON schemas

    2. 2.Any valid JSON schema, including regex patterns and anyOf/oneOf composition, can be passed to Foundation Model APIs.Why is that wrong?

      Only a subset is supported. Regex pattern, composition keywords such as anyOf, oneOf, allOf, prefixItems and $ref, and most type lists are rejected, so keep the schema flat and simple.

      Covered in Writing the output contract: response_format and JSON schemas

    3. 3.Switching model providers means rewriting the structured-output request in that provider's native format.Why is that wrong?

      Databricks uses one response_format field for every supported chat model and handles the provider-specific translation itself.

      Covered in Writing the output contract: response_format and JSON schemas

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “What is the prediction target, and what class of ML problem does that imply”
      ↩︎ From a business goal to inputs, outputs, and constraints
      “What are the serving and production requirements: latency, throughput, and data freshness?”
      ↩︎ From a business goal to inputs, outputs, and constraints
    2. 2.
      “RAG architecture can work with either unstructured or structured supporting data.”
      ↩︎ From a business goal to inputs, outputs, and constraints
    3. 4.
      “Extracting data from large amounts of documents. For example, identifying and classifying product review feedback as negative, positive or neutral.”
      ↩︎ Matching the requirement to a model task
      “Specify your structured outputs using response_format in your chat request.”
      ↩︎ Writing the output contract: response_format and JSON schemas
      “You can choose to generate text, unstructured JSON objects, and JSON objects that adhere to a specific JSON schema.”
      ↩︎ Writing the output contract: response_format and JSON schemas
      “using a simpler JSON schema for JSON schema definitions results in higher quality JSON generation”
      ↩︎ Writing the output contract: response_format and JSON schemas
      “Structured outputs provide a way to generate structured data in the form of JSON objects from your input data.”
      ↩︎ Key concept
      “The following is an example of JSON extraction, but the JSON schema is not known before hand.”
      ↩︎ Exam trap 1
      “Regular expressions using pattern.”
      ↩︎ Exam trap 2
      “You use the same request format regardless of the underlying model provider.”
      ↩︎ Exam trap 3
      “You use the same request format regardless of the underlying model provider.”
      ↩︎ Prediction
    4. 5.
      “Assign predefined categories to documents or text, supporting up to 500+ labels.”
      ↩︎ Matching the requirement to a model task
      “operate directly on the parsed outputs”
      ↩︎ Matching the requirement to a model task
      “Pull structured fields from documents or plain text using a schema you define.”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Choosing Chains, Agents, Tools and Agent Bricks for a GenAI Pipeline

    Spotted a mistake, or was something unclear? Tell us.