What you will be able to do
- Choose between structured outputs (response_format) and function calling for a required output shape
- Configure tool_choice and write tool descriptions so a model calls the right tool
- Pick LLM + prompt, a deterministic chain, a single agent or a multi-agent supervisor, including the Agent Bricks options, for a business use case
1.Making the output machine-readable with response_format
When the next consumer of a model's output is code rather than a person, the output has to have a reliable shape. Prompt instructions are one way to ask for that. Structured outputs enforce it through a response_format field on the chat request. The field works the same way for every supported chat model, and Databricks handles the translation for each provider. Structured outputs are available on Foundation Model APIs pay-per-token and provisioned throughput endpoints. Databricks recommends them for extracting data from large numbers of documents (for example, classifying review feedback as negative, positive or neutral), for batch jobs that need a set format, and for turning unstructured data into structured data.
The schema is part of the request itself, which is why response_format is the Databricks mechanism for fixing the output shape. A step added after the LLM call is only post-processing. In a RAG chain that step applies business logic, adds citations or refines the generated text, so it is not how the shape of the response is set.
The sample below combines a system prompt that defines the model's role with a response_format (defined earlier in the same doc sample) that pins the output to a research_paper_extraction JSON schema.
messages = [{
"role": "system",
"content": "You are an expert at structured data extraction. You will be given unstructured text from a research paper and should convert it into the given structure."
},
{
"role": "user",
"content": "..."
}]
response = client.chat.completions.create(
model="databricks-gpt-oss-20b",
messages=messages,
response_format=response_format
)| Desired output | Setting | Notes |
|---|---|---|
| JSON matching a known schema | "type": "json_schema" with "strict": True | At most 64 keys; pattern, anyOf, oneOf, allOf, prefixItems and $ref are not supported; flatten heavily nested schemas |
| JSON whose schema isn't known in advance | "type": "json_object" | Not supported on Anthropic Claude models |
| Unconstrained text | Omit response_format | The way to get free-form output on Claude |
Checkpoint 1 of 6· Check yourself
You call databricks-claude-sonnet-4-5 with response_format {"type": "json_object"} and stream set to true. What is wrong with this request?
Claude models have extra limits: json_object isn't accepted, and you must set stream to false whenever you pass a response_format.
“Only the json_schema structured output type is supported. json_object is not supported.”Source: docs.databricks.com
Checkpoint 2 of 6· Exam question
A generative AI application must return LLM responses as a JSON object with fixed fields (`sentiment`, `confidence`, `key_phrases`) so a downstream service can parse the result without regex post-processing. Which chain component should the team add after the LLM call to guarantee this structure?
Correct answer: A — An output parser that validates and coerces the LLM response into the required JSON schema
- A. An output parser sits after the LLM call and validates or coerces the raw response into the target schema, guaranteeing the `sentiment`, `confidence`, and `key_phrases` fields the downstream service expects.
- B. A retriever supplies context before generation and cannot enforce or validate the structure of the LLM's output after it is produced.
- C. Instructing the model to respond in JSON without defining a concrete schema still leaves the exact field names and types unenforced, so malformed or inconsistent output can still occur.
- D. Lowering temperature can make wording more consistent but does not guarantee valid, parseable JSON with the exact required fields, so it does not remove the need for schema enforcement.
2.Tools: when the output is an action
Some business requests need the model to act, for example looking up an order, sending an email or querying an index, not just return text. Function calling describes the available functions to the model as JSON-schema tools. The model doesn't run them. It returns a JSON object of arguments, your code makes the call, and the result goes back to the model so it can write the final answer. Use cases include assistants that call other APIs, such as send_email(to: string, body: string), and mapping a phrase like "Who are my top customers?" to a call such as get_customers(min_revenue: int, created_before: string, limit: int). Function definitions use the same restricted JSON schema subset as structured outputs, but the key limit is lower: 16 keys instead of 64.
Checkpoint 3 of 6· Put it in order
Put the function-calling sequence on Databricks in order.
- 1.The model decides whether to call a function and returns JSON arguments that follow your schema
- 2.Your code parses the JSON and calls the function with those arguments
- 3.Call the model with the query and a set of functions in the tools parameter
- 4.Call the model again with the function result appended, so it can summarize the answer for the user
The model only proposes the call. Your code runs it, and a second model call turns the result into a response for the user.
“The model decides whether or not to call the defined functions.”Source: docs.databricks.com
| tool_choice value | Behaviour |
|---|---|
| "auto" (default) | The model decides which functions to call, if any |
| "required" | The model always calls one or more functions and picks which ones |
| {"type": "function", "function": {"name": "my_function"}} | The model calls only that specific function |
| "none" | Function calling is turned off; the model only writes a message for the user |
A retriever can be a tool too. In an agent, you don't hard-wire retrieval as a chain step. You wrap the vector index in a VectorSearchRetrieverTool from Databricks AI Bridge (databricks-langchain or databricks-openai). The model reads tool_name and tool_description to decide when to call the tool, so a vague description leads to the tool being called at the wrong times or not at all.
vs_tool = VectorSearchRetrieverTool(
index_name="catalog.schema.my_databricks_docs_index",
tool_name="databricks_docs_retriever",
tool_description="Retrieves information about Databricks products from official Databricks documentation."
)Checkpoint 4 of 6· Fill the gap
Which LangChain method gives the LLM access to the retriever tool so it can decide to call it?
# Bind the retriever tool to your Langchain LLM of choice
llm = ChatDatabricks(endpoint="databricks-claude-sonnet-4-5")
llm_with_tools = llm. ? ([vs_tool])bind_tools attaches the tool definitions to the chat model. After that, invoke can produce tool calls against the index.
Source: docs.databricks.com3.How much should the LLM orchestrate? Chains, agents and supervisors
Once you know the components, the last decision is who sequences them. Databricks describes a range of options and advises you to start simple, adding agentic behaviour only when you actually need it.
| Design pattern | When to use | Main cost |
|---|---|---|
| LLM + prompt | Generic Q&A; a quick prototype for short-term use | Minimal customization |
| Deterministic chain | Well-defined tasks; static pipelines such as basic RAG; no decisions on the fly | Inflexible; adapting it means changing code |
| Single-agent system | Moderate to complex queries in one domain; some dynamic decisions | Less predictable; must guard against repeated or incorrect tool calls |
| Multi-agent system | Large or cross-functional domains; several expert agents | Complex to orchestrate; harder to trace and debug |
A single agent gets the request and any conversation history, decides whether to call tools, can loop over LLM and tool calls until it reaches its goal, and returns one combined response. Databricks calls it often the sweet spot for enterprise use cases. Because it loops, set iteration limits or timeouts. In a help-desk example, the agent answers a policy question directly but calls lookup_order(customer_id, order_id) for order status, retrying or asking the user for the ID if the tool returns "invalid order number".
In a multi-agent system, each specialized agent has its own expertise, context and tools, and a coordinator or supervisor sends each request to the right agent. That supervisor can be another LLM or a rule-based router. Databricks offers managed versions of these patterns. Knowledge Assistant builds and optimizes domain-specific chatbots through a UI. Supervisor Agent orchestrates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers and custom agents; the Supervisor API docs call it the "Agent Bricks Supervisor Agent". When you need full control, custom agents authored in Python support LangGraph, LangChain, OpenAI and LlamaIndex. Custom agents aren't tied to any one pattern, so you can start with a chain and add autonomy later.
Checkpoint 5 of 6· Check yourself
One assistant must handle finance reconciliations, DevOps incident lookups and marketing campaign questions, each with its own tools and context. Which design fits best?
Very different sub-domains overload a single agent. Separate expert agents behind a supervisor keep each one's tools and context focused.
“If your application spans radically different sub-domains (finance, devops, marketing, etc.), a single agent can become unwieldy”Source: docs.databricks.com
Checkpoint 6 of 6· Exam question
An agent needs to answer questions about current inventory counts that change every few minutes; a static vector search index over product documentation would quickly become stale. Which chain component should the team add so the agent can fetch this real-time data on demand?
Correct answer: A — A Unity Catalog function tool, via `UCFunctionToolkit`, that queries the live inventory table
- A. Wrapping a Unity Catalog function with `UCFunctionToolkit` exposes it as a callable tool, letting the agent query the live inventory table on demand and get up-to-the-minute counts rather than a stale snapshot.
- B. A retriever over an index that is only re-embedded once per day would still return outdated counts for data that changes every few minutes, so it does not meet the real-time requirement.
- C. Embedding yesterday's counts as few-shot examples in the prompt template hardcodes stale numbers and cannot reflect live changes to inventory.
- D. An output parser only reformats text the model has already generated; it cannot supply fresh inventory data if the model was never given access to it in the first place.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Function calling is the best way to turn large volumes of unstructured text into structured records.Why is that wrong?
For batch inference and converting unstructured data into structured data, Databricks recommends structured outputs (response_format). Function calling is for having the model trigger APIs and actions.
Covered in Making the output machine-readable with response_format
2.On a Claude endpoint you can send both tools and a response_format, so the agent calls tools and still returns schema-valid JSON.Why is that wrong?
For Claude models, response_format can't be combined with tools or tool_choice. Split the work into separate calls.
Covered in Tools: when the output is an action
3.A multi-agent supervisor is the safest default for any production assistant.Why is that wrong?
Databricks advises starting simple and adding agentic complexity only when needed. Chains or a single agent are cheaper to run and easier to debug.
Covered in How much should the LLM orchestrate? Chains, agents and supervisors
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“You can choose to generate text, unstructured JSON objects, and JSON objects that adhere to a specific JSON schema.”
↩︎ Making the output machine-readable with response_format“Heavily nested JSON schemas result in lower quality generation.”
↩︎ Making the output machine-readable with response_format“Structured outputs are not supported with streaming.”
↩︎ Making the output machine-readable with response_format“The response_format parameter for Claude structured outputs cannot be combined with tools or tool_choice.”
↩︎ Exam trap 2“Only the json_schema structured output type is supported. json_object is not supported.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-inference-chain-ragOfficial docs
“The LLM's response may be processed further to apply additional business logic, add citations”
↩︎ Making the output machine-readable with response_format - 3.
“The LLM itself does not call these functions, but instead it creates a JSON object”
↩︎ Tools: when the output is an action“The maximum number of keys specified in the JSON schema is 16.”
↩︎ Tools: when the output is an action“For batch inference or data processing tasks, like converting unstructured data into structured data. Databricks recommends using structured outputs.”
↩︎ Exam trap 1“For batch inference or data processing tasks, like converting unstructured data into structured data. Databricks recommends using structured outputs.”
↩︎ Prediction“The model decides whether or not to call the defined functions.”
↩︎ Checkpoint - 4.
“Provide a descriptive tool_description to help the agent understand the tool and determine when to invoke it.”
↩︎ Tools: when the output is an action - 5.
“Infinite loops can occur in any tool-calling scenario, so set iteration limits or timeouts.”
↩︎ How much should the LLM orchestrate? Chains, agents and supervisors“Each agent has its own domain or task expertise, context, and potentially distinct tool sets.”
↩︎ How much should the LLM orchestrate? Chains, agents and supervisors“The supervisor can be another LLM or a rule-based router.”
↩︎ How much should the LLM orchestrate? Chains, agents and supervisors“When building any AI-powered application, start simple.”
↩︎ Exam trap 3“If your application spans radically different sub-domains (finance, devops, marketing, etc.), a single agent can become unwieldy”
↩︎ Checkpoint - 6.
“Build and optimize domain-specific AI chatbots using an intuitive interface.”
↩︎ How much should the LLM orchestrate? Chains, agents and supervisors“Build a supervisor agent that orchestrates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers, and custom agents.”
↩︎ How much should the LLM orchestrate? Chains, agents and supervisors“Supports agents written with any authoring library, including LangGraph, LangChain, OpenAI, and LlamaIndex.”
↩︎ How much should the LLM orchestrate? Chains, agents and supervisors - 7.
“Agent Bricks Supervisor Agent (recommended): Fully declarative with human feedback optimization for highest quality.”
↩︎ How much should the LLM orchestrate? Chains, agents and supervisors