CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 1 · Lesson 3/56

    Structured Outputs, Tools and Agent Patterns for a Desired Output

    Select chain components for a desired model input and output

    10 min read
    1.79% of exam
    7 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Choose between structured outputs (response_format) and function calling for a required output shape
    • Configure tool_choice and write tool descriptions so a model calls the right tool
    • Pick LLM + prompt, a deterministic chain, a single agent or a multi-agent supervisor, including the Agent Bricks options, for a business use case

    1.Making the output machine-readable with response_format

    When the next consumer of a model's output is code rather than a person, the output has to have a reliable shape. Prompt instructions are one way to ask for that. Structured outputs enforce it through a response_format field on the chat request. The field works the same way for every supported chat model, and Databricks handles the translation for each provider. Structured outputs are available on Foundation Model APIs pay-per-token and provisioned throughput endpoints. Databricks recommends them for extracting data from large numbers of documents (for example, classifying review feedback as negative, positive or neutral), for batch jobs that need a set format, and for turning unstructured data into structured data.

    The schema is part of the request itself, which is why response_format is the Databricks mechanism for fixing the output shape. A step added after the LLM call is only post-processing. In a RAG chain that step applies business logic, adds citations or refines the generated text, so it is not how the shape of the response is set.

    The sample below combines a system prompt that defines the model's role with a response_format (defined earlier in the same doc sample) that pins the output to a research_paper_extraction JSON schema.

    A system prompt plus response_format: the input says what to do, the schema fixes the output shapepython
    messages = [{
            "role": "system",
            "content": "You are an expert at structured data extraction. You will be given unstructured text from a research paper and should convert it into the given structure."
          },
          {
            "role": "user",
            "content": "..."
          }]
    
    response = client.chat.completions.create(
        model="databricks-gpt-oss-20b",
        messages=messages,
        response_format=response_format
    )
    Choosing a response_format setting for the output you need
    Desired outputSettingNotes
    JSON matching a known schema"type": "json_schema" with "strict": TrueAt most 64 keys; pattern, anyOf, oneOf, allOf, prefixItems and $ref are not supported; flatten heavily nested schemas
    JSON whose schema isn't known in advance"type": "json_object"Not supported on Anthropic Claude models
    Unconstrained textOmit response_formatThe way to get free-form output on Claude

    Checkpoint 1 of 6· Check yourself

    You call databricks-claude-sonnet-4-5 with response_format {"type": "json_object"} and stream set to true. What is wrong with this request?

    Checkpoint 2 of 6· Exam question

    A generative AI application must return LLM responses as a JSON object with fixed fields (`sentiment`, `confidence`, `key_phrases`) so a downstream service can parse the result without regex post-processing. Which chain component should the team add after the LLM call to guarantee this structure?

    Sources12

    2.Tools: when the output is an action

    Some business requests need the model to act, for example looking up an order, sending an email or querying an index, not just return text. Function calling describes the available functions to the model as JSON-schema tools. The model doesn't run them. It returns a JSON object of arguments, your code makes the call, and the result goes back to the model so it can write the final answer. Use cases include assistants that call other APIs, such as send_email(to: string, body: string), and mapping a phrase like "Who are my top customers?" to a call such as get_customers(min_revenue: int, created_before: string, limit: int). Function definitions use the same restricted JSON schema subset as structured outputs, but the key limit is lower: 16 keys instead of 64.

    Checkpoint 3 of 6· Put it in order

    Put the function-calling sequence on Databricks in order.

    1. 1.The model decides whether to call a function and returns JSON arguments that follow your schema
    2. 2.Your code parses the JSON and calls the function with those arguments
    3. 3.Call the model with the query and a set of functions in the tools parameter
    4. 4.Call the model again with the function result appended, so it can summarize the answer for the user
    tool_choice settings and what the model will do
    tool_choice valueBehaviour
    "auto" (default)The model decides which functions to call, if any
    "required"The model always calls one or more functions and picks which ones
    {"type": "function", "function": {"name": "my_function"}}The model calls only that specific function
    "none"Function calling is turned off; the model only writes a message for the user

    A retriever can be a tool too. In an agent, you don't hard-wire retrieval as a chain step. You wrap the vector index in a VectorSearchRetrieverTool from Databricks AI Bridge (databricks-langchain or databricks-openai). The model reads tool_name and tool_description to decide when to call the tool, so a vague description leads to the tool being called at the wrong times or not at all.

    A retriever exposed as a tool; the description is what the LLM uses to decide when to call itpython
    vs_tool = VectorSearchRetrieverTool(
      index_name="catalog.schema.my_databricks_docs_index",
      tool_name="databricks_docs_retriever",
      tool_description="Retrieves information about Databricks products from official Databricks documentation."
    )

    Checkpoint 4 of 6· Fill the gap

    Which LangChain method gives the LLM access to the retriever tool so it can decide to call it?

    # Bind the retriever tool to your Langchain LLM of choice
    llm = ChatDatabricks(endpoint="databricks-claude-sonnet-4-5")
    llm_with_tools = llm. ? ([vs_tool])

    Sources34

    3.How much should the LLM orchestrate? Chains, agents and supervisors

    Once you know the components, the last decision is who sequences them. Databricks describes a range of options and advises you to start simple, adding agentic behaviour only when you actually need it.

    Agent system design patterns and when each fits
    Design patternWhen to useMain cost
    LLM + promptGeneric Q&A; a quick prototype for short-term useMinimal customization
    Deterministic chainWell-defined tasks; static pipelines such as basic RAG; no decisions on the flyInflexible; adapting it means changing code
    Single-agent systemModerate to complex queries in one domain; some dynamic decisionsLess predictable; must guard against repeated or incorrect tool calls
    Multi-agent systemLarge or cross-functional domains; several expert agentsComplex to orchestrate; harder to trace and debug

    A single agent gets the request and any conversation history, decides whether to call tools, can loop over LLM and tool calls until it reaches its goal, and returns one combined response. Databricks calls it often the sweet spot for enterprise use cases. Because it loops, set iteration limits or timeouts. In a help-desk example, the agent answers a policy question directly but calls lookup_order(customer_id, order_id) for order status, retrying or asking the user for the ID if the tool returns "invalid order number".

    In a multi-agent system, each specialized agent has its own expertise, context and tools, and a coordinator or supervisor sends each request to the right agent. That supervisor can be another LLM or a rule-based router. Databricks offers managed versions of these patterns. Knowledge Assistant builds and optimizes domain-specific chatbots through a UI. Supervisor Agent orchestrates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers and custom agents; the Supervisor API docs call it the "Agent Bricks Supervisor Agent". When you need full control, custom agents authored in Python support LangGraph, LangChain, OpenAI and LlamaIndex. Custom agents aren't tied to any one pattern, so you can start with a chain and add autonomy later.

    Checkpoint 5 of 6· Check yourself

    One assistant must handle finance reconciliations, DevOps incident lookups and marketing campaign questions, each with its own tools and context. Which design fits best?

    Checkpoint 6 of 6· Exam question

    An agent needs to answer questions about current inventory counts that change every few minutes; a static vector search index over product documentation would quickly become stale. Which chain component should the team add so the agent can fetch this real-time data on demand?

    Sources567

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Function calling is the best way to turn large volumes of unstructured text into structured records.Why is that wrong?

      For batch inference and converting unstructured data into structured data, Databricks recommends structured outputs (response_format). Function calling is for having the model trigger APIs and actions.

      Covered in Making the output machine-readable with response_format

    2. 2.On a Claude endpoint you can send both tools and a response_format, so the agent calls tools and still returns schema-valid JSON.Why is that wrong?

      For Claude models, response_format can't be combined with tools or tool_choice. Split the work into separate calls.

      Covered in Tools: when the output is an action

    3. 3.A multi-agent supervisor is the safest default for any production assistant.Why is that wrong?

      Databricks advises starting simple and adding agentic complexity only when needed. Chains or a single agent are cheaper to run and easier to debug.

      Covered in How much should the LLM orchestrate? Chains, agents and supervisors

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “You can choose to generate text, unstructured JSON objects, and JSON objects that adhere to a specific JSON schema.”
      ↩︎ Making the output machine-readable with response_format
      “Heavily nested JSON schemas result in lower quality generation.”
      ↩︎ Making the output machine-readable with response_format
      “Structured outputs are not supported with streaming.”
      ↩︎ Making the output machine-readable with response_format
      “The response_format parameter for Claude structured outputs cannot be combined with tools or tool_choice.”
      ↩︎ Exam trap 2
      “Only the json_schema structured output type is supported. json_object is not supported.”
      ↩︎ Checkpoint
    2. 2.
      “The LLM's response may be processed further to apply additional business logic, add citations”
      ↩︎ Making the output machine-readable with response_format
    3. 3.
      “The LLM itself does not call these functions, but instead it creates a JSON object”
      ↩︎ Tools: when the output is an action
      “The maximum number of keys specified in the JSON schema is 16.”
      ↩︎ Tools: when the output is an action
      “For batch inference or data processing tasks, like converting unstructured data into structured data. Databricks recommends using structured outputs.”
      ↩︎ Exam trap 1
      “For batch inference or data processing tasks, like converting unstructured data into structured data. Databricks recommends using structured outputs.”
      ↩︎ Prediction
      “The model decides whether or not to call the defined functions.”
      ↩︎ Checkpoint
    4. 4.
      “Provide a descriptive tool_description to help the agent understand the tool and determine when to invoke it.”
      ↩︎ Tools: when the output is an action
    5. 5.
      “Infinite loops can occur in any tool-calling scenario, so set iteration limits or timeouts.”
      ↩︎ How much should the LLM orchestrate? Chains, agents and supervisors
      “Each agent has its own domain or task expertise, context, and potentially distinct tool sets.”
      ↩︎ How much should the LLM orchestrate? Chains, agents and supervisors
      “The supervisor can be another LLM or a rule-based router.”
      ↩︎ How much should the LLM orchestrate? Chains, agents and supervisors
      “When building any AI-powered application, start simple.”
      ↩︎ Exam trap 3
      “If your application spans radically different sub-domains (finance, devops, marketing, etc.), a single agent can become unwieldy”
      ↩︎ Checkpoint
    6. 6.
      “Build and optimize domain-specific AI chatbots using an intuitive interface.”
      ↩︎ How much should the LLM orchestrate? Chains, agents and supervisors
      “Build a supervisor agent that orchestrates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers, and custom agents.”
      ↩︎ How much should the LLM orchestrate? Chains, agents and supervisors
      “Supports agents written with any authoring library, including LangGraph, LangChain, OpenAI, and LlamaIndex.”
      ↩︎ How much should the LLM orchestrate? Chains, agents and supervisors
    7. 7.
      “Agent Bricks Supervisor Agent (recommended): Fully declarative with human feedback optimization for highest quality.”
      ↩︎ How much should the LLM orchestrate? Chains, agents and supervisors

    Ready to test yourself?

    Practise the 6 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.