What you will be able to do
- List the steps of a RAG chain in inference order and say which ones are optional
- Explain the design decisions inside each step: query rewriting, embedding model, chunk count, post-processing and guardrails
- Decide when a deterministic chain is the right pattern and when an agent is
Key concept
RAG chain — A RAG chain is the fixed series of steps that runs each time a request comes in. It turns a user query into a retrieval query, fetches supporting chunks, adds them to a prompt and passes that prompt to an LLM. If you are asked to code a simple chain, you are being asked to put these steps together in code.
1.The steps of a RAG chain, in order
A RAG application has two separate parts. The data pipeline runs ahead of time: it pre-processes and indexes documents so they can be retrieved quickly. The RAG chain runs at inference time, once for every request. When an exam requirement says "code a simple chain", it means this inference-time part. The chain takes the user's question, finds supporting data, adds that data to the prompt, and generates an answer from an LLM.
The Databricks reference flow has five steps. The optional preprocessing step changes the raw question into something that searches better: a templated query, a rewritten request, or extracted keywords. Its output is a retrieval query. Retrieval embeds that query and returns the most similar chunks. Augmentation places the chunks and the user's question into a prompt template that tells the model how to use each part. Generation sends that prompt to the LLM, which writes an answer based on the retrieved context. The optional post-processing step then changes the LLM's output. The table below summarises the steps. Each row is one stage, or one function, in your code.
| Step | Required? | Input → output |
|---|---|---|
| User query preprocessing | Optional | User query → retrieval query (templated, rewritten, or keywords) |
| Retrieval | Core | Retrieval query → top (most similar) chunks from the vector database |
| Prompt augmentation | Core | User query + retrieved context → prompt built from a template |
| LLM generation | Core | Augmented prompt → response grounded in the context |
| Post-processing | Optional | LLM response → response with business logic or citations applied |
Checkpoint 1 of 4· Put it in order
Put the steps of a full RAG chain in the order they run for a single request
- 1.Prompt augmentation with retrieved context
- 2.LLM generation
- 3.Retrieval from the vector database
- 4.Post-processing of the response
- 5.User query preprocessing
Preprocessing produces the retrieval query. Retrieval uses that query. The retrieved chunks are added to the prompt, the LLM generates an answer from that prompt, and post-processing changes the LLM's output.
“The output of this step is a retrieval query which will be used in the subsequent retrieval step.”Source: docs.databricks.com
2.Requirements that shape each step
Knowing the order of the steps is not enough. Each step also involves a design decision, and exam requirements are often written to test those decisions.
Query understanding. You can pass the user's query straight to retrieval, and for some queries that works. Databricks says that rewriting the query first is generally beneficial, for example by paraphrasing the conversation history, fixing spelling mistakes, or adding synonyms. Databricks also warns that query rewriting must be done in conjunction with changes to the retrieval component, so plan to change retrieval together with the rewrite rather than adding the rewrite alone.
Retrieval. The query must be embedded using the same embedding model that embedded the document chunks during data preparation. Databricks states this requirement directly; the reason is that retrieval compares the query embedding with the chunk embeddings, and that comparison is designed for embeddings produced by one model. How many chunks you retrieve, and how you combine them with the query in the augmentation step, both have a large effect on answer quality.
Checkpoint 2 of 4· Check yourself
A team indexed their documents with one embedding model. In the chain, they want to embed user queries with a newer and larger model. What happens?
Databricks says the query is embedded with the same model that embedded the document chunks. Matching dimensions is not the stated requirement: the requirement is the same model, so a different model is the wrong choice even if its vectors have the same length.
“the retrieval query is translated into an embedding using the same embedding model that was used to embed the document chunks during data preparation”Source: docs.databricks.com
Post-processing and guardrails. Some requirements cannot be met by the LLM alone: adding citations, applying business rules, checking who the user is. These go around the LLM call. Guardrails can apply anywhere in the chain. They can filter out inappropriate requests, check user permissions before data sources are accessed, and apply content moderation to the generated response. The retrieval step can also be built to return personal or proprietary information only when the user's credentials allow it.
Post-processing. It runs after LLM generation, when your code has both the response and the chunks that were retrieved, so it can add citations according to fixed rules. The model does not have to remember to add them.
3.Why a simple RAG chain is deterministic
Databricks describes several application designs, ordered from least to most autonomous. A basic RAG chain is the standard example of a deterministic chain. The developer decides which tools or models are called, in what order, and with which parameters. The LLM makes no decisions about tool calls. Every request follows the same route: retrieve the top-k chunks, augment the prompt, generate a response.
| Pattern | When to use | Pros | Cons |
|---|---|---|---|
| LLM + prompt | Generic Q&A; quick prototype | Very simple | Minimal customization |
| Deterministic chain | Well-defined tasks; static pipelines such as basic RAG | Simple; easy to audit | Inflexible; requires code changes to adapt |
| Single-agent system | Moderate to complex queries in one domain | Flexible; good default | Less predictable; must guard against bad tool calls |
| Multi-agent system | Large or cross-functional domains | Highly modular | Complex to orchestrate; harder to debug |
Choose a chain when consistency and auditability matter most, and when you want low latency because no extra LLM calls are spent on orchestration decisions. The cost is flexibility. A chain handles unexpected requests poorly, becomes harder to maintain as you add branches, and may need significant refactoring to gain new capabilities. Databricks advises starting simple and adding agent-like behaviour only when you actually need it.
Checkpoint 3 of 4· Match them up
Match each design pattern to the situation it fits
Tap a term, then the definition that fits it.
The design-patterns table lists basic RAG as the example use case for a deterministic chain. Use an agent when the model needs to decide which tools to call.
“Static pipelines such as basic RAG”Source: docs.databricks.com
A short preview before the next question, which draws on the packaging step covered in detail on the next page. When chain code is packaged as an MLflow pyfunc model, Databricks says it needs two functions: load_context, for anything loaded once so the model can operate, and predict, which holds the logic that runs on every request. Pre- and post-processing of inputs and outputs belong in that custom code around the model call.
Checkpoint 4 of 4· Exam question
A generative AI engineer is coding a chain as an MLflow pyfunc model. The chain must strip and normalize incoming user queries before they reach the underlying LLM, and reformat the LLM's raw output into a fixed JSON schema before returning a response. Where should the engineer place this pre-processing and post-processing logic?
Correct answer: A — Inside the `predict` method, so normalization runs before the LLM call and reformatting runs on its output.
- A. Correct: `predict` runs on every incoming request in a pyfunc model, making it the right place to wrap the LLM call with input normalization beforehand and output reformatting afterward.
- B. Incorrect: `load_context` executes once when the model is loaded, to initialize artifacts and dependencies, not to transform each request's query text or each response's content.
- C. Incorrect: a model signature only declares the expected input and output schema types for validation; it has no mechanism to execute normalization or reformatting code.
- D. Incorrect: splitting transformation logic into a separate script outside the pyfunc model breaks the self-contained chain, and Model Serving does not chain a logged model to an external post-processing script this way.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.In a simple RAG chain, the LLM decides when to call the retriever.Why is that wrong?
In a deterministic chain, the developer fixes the tools, their order and their parameters. If the model chooses which tools to call, you are using the single-agent pattern.
Covered in Why a simple RAG chain is deterministic
2.Any good embedding model can embed the user's query at inference time.Why is that wrong?
The query must be embedded with the same model that embedded the document chunks. Otherwise the similarity comparison does not work.
Covered in Requirements that shape each step
3.You can add query rewriting to a chain on its own, without changing anything else.Why is that wrong?
Databricks warns that query rewriting must be done together with changes to the retrieval component.
Covered in Requirements that shape each step
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“RAG chain (Retrieval, Augmentation, Generation): Call a series (or chain) of steps to:”
↩︎ The steps of a RAG chain, in order“The retrieval step can be designed to selectively retrieve personal or proprietary information based on user credentials.”
↩︎ Requirements that shape each step - 2.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-inference-chain-ragOfficial docs
“The prompt that will be sent to the LLM is formed by augmenting the user's query with the retrieved context”
↩︎ The steps of a RAG chain, in order“determining how many chunks to retrieve in step 2 and how to combine them with the user's query in step 3 can significantly impact”
↩︎ Requirements that shape each step“checking user permissions before accessing data sources, and applying content moderation techniques to the generated responses”
↩︎ Requirements that shape each step“The series, or chain of steps that are invoked at inference time is commonly referred to as the RAG chain.”
↩︎ Key concept“the retrieval query is translated into an embedding using the same embedding model that was used to embed the document chunks during data preparation”
↩︎ Exam trap 2“(Optional) User query preprocessing: In some cases, the user's query is preprocessed to make it more suitable for querying the vector database.”
↩︎ Prediction“The output of this step is a retrieval query which will be used in the subsequent retrieval step.”
↩︎ Checkpoint“the retrieval query is translated into an embedding using the same embedding model that was used to embed the document chunks during data preparation”
↩︎ Checkpoint - 3.
“it is generally beneficial to reformulate the query before the retrieval step.”
↩︎ Requirements that shape each step“Query rewriting must be done in conjunction with changes to the retrieval component”
↩︎ Requirements that shape each step“Query rewriting must be done in conjunction with changes to the retrieval component”
↩︎ Exam trap 3 - 4.
“the developer defines which tools or models are called, in what order, and with which parameters.”
↩︎ Why a simple RAG chain is deterministic“When you want to minimize latency by avoiding multiple LLM calls for orchestration decisions.”
↩︎ Why a simple RAG chain is deterministic“Can require significant refactoring to accommodate new capabilities.”
↩︎ Why a simple RAG chain is deterministic“The LLM does not make decisions about which tools to call or in what order.”
↩︎ Exam trap 1“Static pipelines such as basic RAG”
↩︎ Checkpoint - 5.https://docs.databricks.com/aws/en/machine-learning/model-serving/deploy-custom-python-codeOfficial docs
“predict - this function houses all the logic that is run every time an input request is made.”
↩︎ Why a simple RAG chain is deterministic“Your application requires the model's raw outputs to be post-processed for consumption.”
↩︎ Why a simple RAG chain is deterministic