What you will be able to do
- Put the stages of a RAG chain in order and say what each one takes in and passes on
- Choose query-understanding steps (rewriting, filter extraction, intent classification, entity extraction) to fit a business request
- Choose semantic, keyword or hybrid retrieval based on how users word their queries
Key concept
Deterministic chain — A pipeline where you, the developer, wire a fixed sequence of components (query preprocessing, retriever, prompt template, LLM, post-processor), so every request takes the same path. Picking chain components means deciding which stages your input and output need, and in what order.
1.The chain as an input-to-output pipeline
A business request such as "answer employee questions from our HR policies" has a fixed input (a user's question) and a fixed output (a grounded answer, possibly with citations). The components between them form a chain. Databricks calls the steps that run at inference time the RAG chain. Each step takes something in and passes something on:
1. (Optional) query preprocessing. This step takes the raw user question and returns a *retrieval query*. It might put the question into a template, have another model rewrite it, or pull out keywords. A shorter Databricks summary of the chain calls this first step "understand the user's question". 2. Retrieval. This step takes the retrieval query and returns ranked chunks. The query is embedded with the same embedding model that embedded the documents, and the most similar chunks come back. 3. Prompt augmentation. This step takes the user query plus the chunks and returns a prompt. A *prompt template* combines them and tells the model how to use each part. The template is also where you add instructions that control the response format. 4. LLM generation. This step takes the augmented prompt and returns a grounded response. 5. (Optional) post-processing. This step takes the raw response and returns the final output. It applies business logic, adds citations or applies other rules.
Note the order: the retriever does not feed the LLM directly. Prompt augmentation sits between retrieval and generation, because the LLM needs a finished prompt, not a bare list of chunks.
The two optional stages are the ones you add or drop depending on the use case. If users phrase questions loosely, add preprocessing. If the business needs citations or rule enforcement on the output, add post-processing. Guardrails such as request filtering, permission checks and content moderation can sit at any point along the chain.
Three of these components shape the output you get, so choose them deliberately:
- The prompt template is where you control the response format. If the output must be JSON, Databricks structured outputs let you request a defined JSON format through a response_format field on a chat request. Databricks recommends this for extracting data from documents, batch inference with a required output format, and turning unstructured data into structured data. Structured outputs are not supported with streaming for Claude models, so check the model's constraints before you commit to them.
- The LLM is a component you select, together with its parameters. The goal is to balance performance, latency and cost for your application, so a task that needs fast, cheap answers may call for a different model than one that needs the most careful reasoning.
- Post-processing and guardrails are the output-side components. Use them to keep responses on-topic, factually consistent and within specific guidelines or constraints, and to add citations or business logic.
Checkpoint 1 of 4· Put it in order
Put the core RAG chain steps in the order they run at inference time. The first step corresponds to the optional query preprocessing stage.
- 1.Retrieve supporting data
- 2.Augment the prompt with supporting data
- 3.Generate a response from an LLM using the augmented prompt
- 4.Understand the user's question
Each step consumes the output of the previous one. The retriever needs a query, the template needs the retrieved chunks, and the LLM needs the finished prompt.
“Understand the user's question. Retrieve supporting data. Augment the prompt with supporting data. Generate a response from an LLM using the augmented prompt.”Source: docs.databricks.com
None of them at run time. You decide when you build the chain: every request goes through the same steps, and the LLM never chooses which tools to call or in what order. That is why chains are predictable and easy to audit, but adapting one means changing the code.
Checkpoint 2 of 4· Exam question
A team is building a RAG-based customer support chain on Databricks. The chain must fetch the most relevant product documentation chunks from a Databricks AI Search (Vector Search) index before the LLM generates an answer. Which chain component should the team add immediately before the LLM call to supply this grounding context?
Correct answer: A — A retriever wired to the AI Search index, which returns the top-k relevant chunks as grounding context
- A. A retriever configured against an AI Search index performs semantic search over the embedded documentation and returns the top-k most relevant chunks, which is exactly the grounding step a RAG chain needs before the LLM call.
- B. An output parser runs after the LLM generates a response to enforce a structured format; it plays no role in supplying context before generation, so it does not solve the retrieval requirement here.
- C. Hardcoding the entire documentation set into the system prompt does not scale, quickly exceeds context-window limits, and skips the relevance filtering that a retriever provides.
- D. A Unity Catalog function tool is suited to querying live, structured, transactional data like orders, not for semantic retrieval over unstructured product documentation chunks.
2.Query understanding: matching a model task to the input
The query-understanding stage is where you choose which model task to run on the input, and each task produces a different output for the next step:
- Query rewriting produces one or more better queries. Examples include paraphrasing conversation history in a multi-turn chat, fixing spelling, and swapping in synonyms. A rewritten query only helps if the retriever is changed to use it. - Filter extraction produces structured parameters for the retriever, such as a time period ("reports from 2023"), a product, or a city or country. This changes components on both sides: the data pipeline has to extract the matching metadata for each chunk, and the retriever has to accept and apply the filters. - Intent classification sorts the query into predefined categories, such as "product information", "troubleshooting" or "account management". Its output is a label that later steps can use. - Entity extraction pulls specific values out of the query, such as product names, reported errors or account numbers.
You can do all of this in one carefully written LLM call, or split it across several calls. The Databricks example for a customer support bot uses a sequence: classify the intent, then extract entities based on that intent, then use both to rewrite the query into a more specific, targeted form. So in this design the intent label shapes which entities are extracted, and the entities and intent together shape the final retrieval query. Splitting the work adds latency but gives you finer control. The same rule of thumb applies whenever one prompt would have to carry many complex logic steps.
Checkpoint 3 of 4· Match them up
Match each query-understanding task to the output it produces for the next component.
Tap a term, then the definition that fits it.
Each task turns the free-text input into a different structured output. A multi-step design chains them: intent first, then entities, then the rewrite.
“use another LLM call to extract relevant entities from the query, such as product names, reported errors, or account numbers.”Source: docs.databricks.com
Sources2
3.Choosing the retriever from the shape of the query
The retriever is the component that most directly reflects what your users' input looks like. Ask whether their queries share *concepts* with the documents, share exact *terms* with them, or both. For unstructured data, retrieval uses semantic search, keyword search or a combination of the two, and the choice depends on the data and the queries you expect.
| Strategy | Relevance rule | Example query vs. document | Technical approach |
|---|---|---|---|
| Semantic search | Same concepts in query and document | "how do i turn my phone on?" vs. a manual section called "toggling the power" | Embeddings in a continuous vector space |
| Keyword search | Same words; more shared words means more relevant | "what does model HD7-8D do?" | bag-of-words, TF-IDF, BM25 |
| Hybrid search | Semantic search first, then keyword refinement of that smaller set, then the two scores are combined | "how do I turn on my HD7-8D?" | Re-ranking, e.g. reciprocal rank fusion or a re-ranking model |
Hybrid search is a two-step process rather than two independent searches. It first runs a semantic search to get a set of conceptually relevant documents, then applies keyword search to that reduced set to refine the results on exact matches, and finally combines the scores from both steps to rank the documents.
On top of the core strategy, Databricks lists techniques that you can add to the retrieval stage. Each one needs coordinated changes elsewhere in the chain:
- Query expansion uses several variations of the retrieval query to capture a wider range of relevant documents. The variations are typically generated in the query-understanding component, so that component must change too. - Re-ranking re-orders the initial chunks using extra criteria (for example, sort by time) or a reranker model, so the most relevant chunks come first. - Metadata filtering narrows the search space using attributes such as document type, creation date, author or domain tags, and can be combined with semantic or keyword search. The filters come from the query-understanding step, and the metadata must be extracted in the data pipeline, so both must change along with the retriever.
Checkpoint 4 of 4· Check yourself
An internal parts catalogue bot mostly gets queries like "specs for part ZX-440B", where the exact code matters more than the meaning. Which retrieval strategy fits that input best?
Product codes carry no meaning an embedding could match on. Keyword matching is the strategy built for term-focused queries such as product names.
“Scenarios requiring precise keyword matches, ideal for specific term-focused queries such as product names.”Source: docs.databricks.com
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.In a deterministic RAG chain, the LLM decides when to call the retriever.Why is that wrong?
In a chain, the developer fixes which components run and in what order. Letting the model choose tools makes it a single-agent system.
Covered in The chain as an input-to-output pipeline
2.Adding a query-rewriting or filter-extraction step is a self-contained change at the front of the chain.Why is that wrong?
These components depend on other parts of the chain. Rewriting needs matching changes in the retriever, and filter extraction also needs metadata extracted in the data pipeline.
Covered in Query understanding: matching a model task to the input
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-inference-chain-ragOfficial docs
“The series, or chain of steps that are invoked at inference time is commonly referred to as the RAG chain.”
↩︎ The chain as an input-to-output pipeline“in a template that instructs the model how to use each component, often with additional instructions to control the response format.”
↩︎ The chain as an input-to-output pipeline“The LLM's response may be processed further to apply additional business logic, add citations, or otherwise refine the generated text”
↩︎ The chain as an input-to-output pipeline“using the same embedding model that was used to embed the document chunks during data preparation”
↩︎ The chain as an input-to-output pipeline - 2.
“Selecting the most appropriate model (and model parameters) for your application to optimize/balance performance, latency, and cost.”
↩︎ The chain as an input-to-output pipeline“Applying additional processing steps and safety measures to ensure the LLM-generated responses are on-topic, factually consistent”
↩︎ The chain as an input-to-output pipeline“you might use one LLM call to classify the query intent, another to extract relevant entities, and a third to rewrite the query”
↩︎ Query understanding: matching a model task to the input“Although this approach may add some latency to the overall process, it can allow for more fine-grained control”
↩︎ Query understanding: matching a model task to the input“Query rewriting must be done in conjunction with changes to the retrieval component”
↩︎ Query understanding: matching a model task to the input“Use the extracted intent and entities to rewrite the original query into a more specific and targeted format”
↩︎ Query understanding: matching a model task to the input“depends on the specific requirements of your application, the nature of the data, and the types of queries you expect to handle.”
↩︎ Choosing the retriever from the shape of the query“First, it performs a semantic search to retrieve a set of conceptually relevant documents.”
↩︎ Choosing the retriever from the shape of the query“Query expansion must be done in conjunction with changes to the query understanding component (RAG chain).”
↩︎ Choosing the retriever from the shape of the query“apply additional ranking criteria (for example, sort by time) or a reranker model to re-order the results.”
↩︎ Choosing the retriever from the shape of the query“Metadata filtering must be done in conjunction with changes to the query understanding (RAG chain) and metadata extraction (data pipeline) components.”
↩︎ Choosing the retriever from the shape of the query“Filter extraction must be done in conjunction with changes to both metadata extraction data pipeline and retriever chain components.”
↩︎ Exam trap 2“it is generally beneficial to reformulate the query before the retrieval step.”
↩︎ Prediction“use another LLM call to extract relevant entities from the query, such as product names, reported errors, or account numbers.”
↩︎ Checkpoint“Scenarios requiring precise keyword matches, ideal for specific term-focused queries such as product names.”
↩︎ Checkpoint - 3.
“Structured outputs on Databricks let you generate responses in a defined JSON format as part of your AI application workflows.”
↩︎ The chain as an input-to-output pipeline“Structured outputs are not supported with streaming.”
↩︎ The chain as an input-to-output pipeline
Also cited
“the developer defines which tools or models are called, in what order, and with which parameters.”
↩︎ Key concept“The LLM does not make decisions about which tools to call or in what order.”
↩︎ Exam trap 1“Understand the user's question. Retrieve supporting data. Augment the prompt with supporting data. Generate a response from an LLM using the augmented prompt.”
↩︎ Checkpoint