What you will be able to do
- Order the components of a RAG chain and say what each one contributes to the input/output contract
- Decide between a deterministic chain and a tool-calling single agent based on the business requirements
- Assign tools by category and choose Knowledge Assistant, Supervisor Agent, Information Extraction, or a custom agent for a given use case
1.Chain components, each with its own input and output
Once the inputs and outputs are specified, the pipeline is the set of components that turns one into the other. RAG is the most common shape when the business needs answers about its own data. It is "especially valuable for answering questions about proprietary, frequently changing, or domain-specific information." At inference time, the RAG chain is a series of steps, and each step hands a defined artefact to the next.
| Step | Takes in | Produces |
|---|---|---|
| User query preprocessing (optional) | Raw user query | A retrieval query (templated, rewritten, or keyword-extracted) |
| Retrieval | Retrieval query, embedded with the same model used for the chunks | Top-ranked similar chunks |
| Prompt augmentation | User query plus retrieved context | A prompt built from a template, often with response-format instructions |
| LLM generation | Augmented prompt | A response grounded in the context |
| Post-processing (optional) | LLM response | Response with business logic applied or citations added |
Two of these steps carry the output contract. The prompt template is where formatting instructions go: retrieved context is combined with the query "in a template that instructs the model how to use each component, often with additional instructions to control the response format". Post-processing is where you enforce rules the model shouldn't be trusted with. The LLM's response "may be processed further to apply additional business logic, add citations, or otherwise refine the generated text". If the business requires every answer to cite its source, that requirement belongs in these two steps. Guardrails such as permission checks and content moderation can also be applied anywhere along the chain.
Checkpoint 1 of 5· Put it in order
Put the RAG inference chain steps in the order they run
- 1.Retrieve the most similar chunks from the vector database
- 2.Augment the prompt template with the query and retrieved context
- 3.Post-process the response, for example by adding citations
- 4.Preprocess the user query into a retrieval query
- 5.Generate a grounded response with the LLM
Each step consumes the previous step's output: the retrieval query drives retrieval, the chunks fill the template, the augmented prompt goes to the LLM, and the response is refined last.
“The output of this step is a retrieval query which will be used in the subsequent retrieval step.”Source: docs.databricks.com
2.Deterministic chain or tool-calling agent?
A fixed chain always runs the same steps. That is the right choice when the business requirement is narrow and predictable. Databricks says to use a deterministic chain "For well-defined tasks with predictable workflows." It is also the right choice when consistency and auditing come first, or when latency must stay low because no LLM calls are spent on orchestration decisions. The cost is flexibility: unexpected requests are handled poorly, and new capabilities can require significant refactoring.
Some goals need branching that a fixed chain can't express. Take the call-centre request "Can you help me return my last order?" The input is a single sentence. The output is a confirmation with a shipping label. In between, the system plans, looks up data, checks a rule, and takes an action: "The agent queries the order database to retrieve the relevant order and references a policy document." Which steps run depends on what each one returns.
In a single-agent system, "The LLM adaptively decides which tools to use, when to make more LLM calls, and when to stop." Its input contract is wider than a chain's. It accepts the user query plus relevant context such as conversation history, and it returns a single cohesive response after combining whatever its tools return. That flexibility brings new duties. Infinite loops can occur in any tool-calling setup, so set iteration limits or timeouts. And if the application spans very different sub-domains (finance, devops, marketing), one agent can become unwieldy. That is the point where a coordinating multi-agent design becomes worth considering.
| Requirement | Deterministic chain | Single-agent system |
|---|---|---|
| Workflow shape | Well-defined, predictable | Varied queries within one cohesive domain |
| Auditability and consistency | Highest | Lower; must guard against invalid or repeated tool calls |
| Latency | Typically lower (fewer orchestration LLM calls) | Can loop through repeated LLM or tool calls |
| Handling unexpected requests | Limited | Adapts by choosing which tools, if any, to call |
Checkpoint 2 of 5· Check yourself
A compliance team needs a nightly job that answers the same templated question about each new policy document. Every run must follow identical, auditable steps with minimal latency. Which design fits?
Predictable workflows where consistency, auditing and low latency come first are the textbook case for a deterministic chain.
“When consistency and auditing are top priorities.”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
A finance team's business goal is an invoice-processing pipeline in which every invoice must be checked against three specific validation rules in a fixed, unchanging order, with no deviation permitted for compliance reasons. Which agent design pattern best matches this input/output requirement?
Correct answer: A — A deterministic chain in which the developer fixes which checks run, in what order, and with what parameters for every invoice
- A. When the exact set and order of checks must never vary for compliance reasons, a deterministic chain is correct because the developer, not the model, fixes the sequence of calls and parameters for every invoice.
- B. Letting the LLM adaptively choose which checks to run and in what order introduces exactly the kind of variability the compliance requirement rules out, since the model could skip or reorder a validation step.
- C. Dynamic routing between specialized agents is meant for domains requiring case-by-case judgment, not for a task that must apply the same fixed three checks in the same order every time.
- D. Free-text descriptions of possible issues do not guarantee that all three rules were actually checked in the required order, so this pattern cannot satisfy the fixed-sequence compliance requirement.
Sources3
3.Choosing tools for the inputs the agent must fetch
An agent's tools follow from the inputs it needs and the actions the business wants taken. "Tools are single-interaction functions that an LLM can invoke to accomplish a well-defined task." The model generates the parameters, and the tool returns a direct result. Databricks groups them into three kinds. Tools that retrieve or analyse data include semantic retrieval over a vector index, structured retrieval through SQL or APIs, web search, classic ML models, and other AI models. Tools that modify an external system include CRM or service API calls and email or messaging notifications. Tools that run logic include sandboxed code execution. Tools can be built into the agent's logic or exposed through standard interfaces such as MCP.
Here the input analysis from scoping pays off. Unstructured supporting data (manuals, policies) points to a semantic retrieval tool. Tabular data (orders, customer records) points to structured retrieval. A requirement such as "and then file the return" points to a tool that changes state. In the help-desk example, a lookup_order(customer_id, order_id) function is that structured tool. If it returns "invalid order number," the agent can retry or ask the user for the correct ID.
Checkpoint 4 of 5· Check yourself
An agent must answer "What were my last three transactions?" from a SQL database. Which tool category supplies that input?
Transaction rows are structured, schema-bound data, so the matching tool runs SQL or calls an API. Semantic retrieval is for unstructured text.
“Structured retrieval: Run SQL queries or use APIs to retrieve structured information.”Source: docs.databricks.com
Sources4
4.Agent Bricks or a custom agent
Databricks offers a range of build options, from guided to fully custom. Several common business goals map directly onto a guided Agent Brick, so check those before writing a chain from scratch. Knowledge Assistant fits when the input is a document collection and the output is cited answers: "Use Knowledge Assistant to create a chatbot that can answer questions about your documents and provide high-quality responses with citations." Its listed use cases are product documentation Q&A, HR policy questions, and customer inquiries answered from support knowledge bases. It produces an agent endpoint for downstream applications.
Supervisor Agent fits when one request may need several specialised sources. It coordinates Genie Agents, agent endpoints, Unity Catalog functions, MCP servers, and custom agents to "work together to complete complex tasks across different, specialized domains". Examples include market analysis across research reports and usage data, and customer service that combines policy, FAQ, and account questions. When the output is structured fields rather than conversation, Information Extraction in the Agent Bricks UI (ai_extract in SQL) pulls fields against your schema. When none of these fit, "Custom agents allows you to build and deploy agents using custom code or third-party agent authoring libraries."
Checkpoint 5 of 5· Match them up
Match each business requirement to the Databricks option that fits it
Tap a term, then the definition that fits it.
Knowledge Assistant covers cited Q&A over documents, Supervisor Agent coordinates multiple specialised agents and tools, Information Extraction returns schema-defined fields, and custom agents cover everything else in code.
“Answer employee questions related to HR policies.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A tool-calling agent is always the better design, because it can handle anything a fixed chain can.Why is that wrong?
Deterministic chains are recommended for predictable workflows where auditability, consistency and low latency matter. Agents add flexibility but also loop risk and less predictability.
Covered in Deterministic chain or tool-calling agent?
Practise it for real
Send a structured-output extraction request to a Foundation Model API chat model and confirm the response follows your output contract.
1.Create an OpenAI client pointed at your Databricks base URL with your Databricks token, as in the structured outputs documentation.
Why: Databricks Foundation Model APIs accept OpenAI-compatible chat requests.
You should see: A client object, with no errors on creation.
2.Define a response_format of type json_schema named research_paper_extraction with title, authors, abstract and keywords properties and strict set to True.
Why: This puts the output contract into the request instead of relying on prompt wording.
You should see: A Python dict that uses only supported schema features (no pattern, anyOf, oneOf or $ref).
3.Build messages with the extraction system prompt and paste a paper description into the user message, then call client.chat.completions.create with model databricks-gpt-oss-20b.
Why: The system prompt sets the task, the user message is the pipeline input, and response_format constrains the output.
You should see: A response whose content holds the extracted fields; for this model, take the block with type text and parse it as JSON.
4.Change only the model value to databricks-claude-sonnet-4-5 and run the request again.
Why: This confirms the same response_format works across providers.
You should see: The same JSON field structure, subject to the documented Claude limitations.
5.Replace response_format with {"type": "json_object"} and a prompt naming the fields to extract.
Why: This shows the mode for when the schema isn't known in advance.
You should see: Valid JSON, with field names guided only by the prompt.
Stuck? Get a nudge
If parsing fails on the gpt-oss model, check whether content came back as a list of blocks rather than a plain string.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“This approach is especially valuable for answering questions about proprietary, frequently changing, or domain-specific information.”
↩︎ Chain components, each with its own input and output - 2.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-inference-chain-ragOfficial docs
“in a template that instructs the model how to use each component, often with additional instructions to control the response format”
↩︎ Chain components, each with its own input and output“The LLM's response may be processed further to apply additional business logic, add citations, or otherwise refine the generated text”
↩︎ Chain components, each with its own input and output“The output of this step is a retrieval query which will be used in the subsequent retrieval step.”
↩︎ Checkpoint - 3.
“The LLM adaptively decides which tools to use, when to make more LLM calls, and when to stop.”
↩︎ Deterministic chain or tool-calling agent?“The agent queries the order database to retrieve the relevant order and references a policy document.”
↩︎ Deterministic chain or tool-calling agent?“Infinite loops can occur in any tool-calling scenario, so set iteration limits or timeouts.”
↩︎ Deterministic chain or tool-calling agent?“For well-defined tasks with predictable workflows.”
↩︎ Exam trap 1“This design pattern is often the sweet spot for enterprise use cases”
↩︎ Prediction“When consistency and auditing are top priorities.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/agents/conceptsOfficial docs
“Tools are single-interaction functions that an LLM can invoke to accomplish a well-defined task.”
↩︎ Choosing tools for the inputs the agent must fetch“Custom agents allows you to build and deploy agents using custom code or third-party agent authoring libraries.”
↩︎ Agent Bricks or a custom agent“Structured retrieval: Run SQL queries or use APIs to retrieve structured information.”
↩︎ Checkpoint - 5.
“Use Knowledge Assistant to create a chatbot that can answer questions about your documents and provide high-quality responses with citations.”
↩︎ Agent Bricks or a custom agent“Answer employee questions related to HR policies.”
↩︎ Checkpoint - 6.
“work together to complete complex tasks across different, specialized domains”
↩︎ Agent Bricks or a custom agent