What you will be able to do
- Compose a prompt, LLM and output parser into a LangChain chain using the pipe operator
- Build the retrieval component with VectorSearchRetrieverTool or index.similarity_search
- Wrap custom pre- and post-processing in an MLflow pyfunc model using load_context and predict
1.Composing a chain with LangChain
LangChain links chain components with the pipe operator. Each component's output becomes the next component's input. There are three parts: a prompt template with named placeholders, a chat model, and a final step that converts the model's message into a plain string. The example below comes from the Databricks MLflow tracing documentation and uses an OpenAI chat model. In a Databricks chain, the llm slot is usually filled by ChatDatabricks(endpoint=...) from the databricks-langchain package, which is covered in the next section.
llm = ChatOpenAI(model="gpt-4o-mini", temperature=0.7, max_tokens=1000)
prompt_template = PromptTemplate.from_template(
"Answer the question as if you are {person}, fully embodying their style, wit, personality, and habits of speech. "
"Emulate their quirks and mannerisms to the best of your ability, embracing their traits—even if they aren't entirely "
"constructive or inoffensive. The question is: {question}"
)Checkpoint 1 of 5· Fill the gap
Which component completes the chain so that it returns a plain string instead of a chat message object?
chain = prompt_template | llm | ? ()
# Let's test another call
chain.invoke(
{
"person": "Linus Torvalds",
"question": "Can I just set everyone's access to sudo to make things easier?",
}
)The order is prompt template, then chat model, then StrOutputParser. The dictionary passed to invoke supplies the template's {person} and {question} placeholders.
Source: docs.databricks.comThe input to invoke is a dictionary whose keys match the template's placeholders. A missing key means the prompt cannot be filled. The same page also warns against hardcoding provider API keys in chain code. For production, use AI Gateway or Databricks secrets.
Checkpoint 2 of 5· Exam question
A generative AI engineer is assembling a simple chain in LangChain on Databricks that takes a user question, applies a `ChatPromptTemplate`, sends it to a `ChatDatabricks` LLM, and parses the output with a `StrOutputParser`. Which code correctly composes these three components into a single runnable chain using LangChain Expression Language?
Correct answer: A — Chain the components with the pipe operator: `prompt | llm | output_parser`, composing them into one Runnable in sequence.
- A. Correct: the pipe operator is LangChain Expression Language's standard syntax for composing Runnables, so linking the prompt template, chat model, and output parser this way builds a single sequential chain.
- B. Incorrect: `LLMChain` is the legacy chain class and is not constructed by passing a prompt, chat model, and parser as three positional arguments in this way.
- C. Incorrect: `SequentialChain` composes multiple named sub-chains with declared input/output keys for multi-step workflows, it is not how a single prompt-LLM-parser pipeline built from these three objects is expressed.
- D. Incorrect: `RunnableSequence` is not instantiated with a `steps` keyword this way; the pipe operator is the documented syntax that builds a `RunnableSequence` under the hood.
Sources1
2.Adding the retrieval step
A RAG chain also needs a retriever. Databricks AI Bridge packages such as databricks-langchain include VectorSearchRetrieverTool, which wraps an existing index, and ChatDatabricks, which calls a Databricks serving endpoint. You can test the tool by calling invoke on it directly.
from databricks_langchain import VectorSearchRetrieverTool, ChatDatabricks
# Initialize the retriever tool.
vs_tool = VectorSearchRetrieverTool(
index_name="catalog.schema.my_databricks_docs_index",
tool_name="databricks_docs_retriever",
tool_description="Retrieves information about Databricks products from official Databricks documentation."
)
# Run a query against the vector search index locally for testing
vs_tool.invoke("Databricks Agent Framework?")
# Bind the retriever tool to your Langchain LLM of choice
llm = ChatDatabricks(endpoint="databricks-claude-sonnet-4-5")
llm_with_tools = llm.bind_tools([vs_tool])The last two lines matter for design. bind_tools lets the LLM decide whether to call the retriever, which is how a single agent works. In a deterministic chain, your own code calls the retriever every time and passes the results into the prompt. If the index uses self-managed embeddings, or is a direct-access index, the tool also needs a text_column and an embedding model, for example DatabricksEmbeddings(endpoint="databricks-bge-large-en").
A deterministic RAG chain follows the same fixed order for every request: retrieve the top-k results from the index, augment a prompt by combining the user request with the retrieved context, then generate a response by sending the augmented prompt to an LLM. In LangChain terms, your code fetches the chunks first, supplies them together with the question as the placeholders of a prompt template, and then pipes that template into the chat model and the output parser, exactly as in the previous section. The LLM never decides whether to retrieve. That fixed flow is what makes a deterministic chain predictable and easy to audit.
You can also skip LangChain and query the index yourself with index.similarity_search. Which argument you pass depends on how the index got its embeddings.
| Argument | Use it when |
|---|---|
| query_text | Delta Sync index with embeddings computed by Databricks |
| query_vector | Delta Sync index with pre-calculated embeddings |
| query_type | Choosing a different search type, such as "hybrid" or "FULL_TEXT" (Beta) |
| columns / num_results | Choosing which fields to return and how many results |
Checkpoint 3 of 5· Fill the gap
The index stores pre-calculated embeddings, so your chain embeds the query itself. Which argument goes in the blank?
# Delta Sync Index with pre-calculated embeddings
results2 = index.similarity_search(
? =[0.9] * 1024,
columns=["id", "text"],
num_results=2
)Use query_text when Databricks computes the embeddings. With pre-calculated embeddings you pass the query's own vector as query_vector.
Source: docs.databricks.com3.Wrapping custom logic as a pyfunc model
If a requirement calls for logic that a framework does not provide, such as reformatting inputs before the model or cleaning raw outputs afterwards, MLflow's pyfunc can package any Python code as a model. The Databricks guidance lists cases where pyfunc fits: the model needs preprocessing, its raw outputs need post-processing, it has branching logic per request, or you want to deploy fully custom code.
class CustomModel(mlflow.pyfunc.PythonModel):
def load_context(self, context):
self.model = torch.load(context.artifacts["model-weights"])
from preprocessing_utils.my_custom_tokenizer import CustomTokenizer
self.tokenizer = CustomTokenizer(context.artifacts["tokenizer_cache"])
def format_inputs(self, model_input):
# insert some code that formats your inputs
pass
def format_outputs(self, outputs):
predictions = (torch.sigmoid(outputs)).data.numpy()
return predictions
def predict(self, context, model_input):
model_input = self.format_inputs(model_input)
outputs = self.model.predict(model_input)
return self.format_outputs(outputs)The two required methods split the work. load_context holds everything that only needs loading once, such as weights and tokenizers, so that each request has less to load. predict runs on every request. In this example it calls the preprocessing helper format_inputs, then the model, then the post-processing helper format_outputs. A chain fits the same shape: the optional query preprocessing step of a RAG chain maps to the input formatting, and the optional post-processing step, such as adding citations or applying business rules to the LLM's response, maps to the output formatting. Before deploying, you can check that the model can be served by using mlflow.models.predict. After logging it, you can register it and serve it from a Model Serving endpoint.
Logging matters for serving. Databricks Runtime ML includes mlflow-skinny by default, and if you log without specifying pip_requirements, MLflow captures mlflow-skinny in the model's conda.yaml. Model Serving requires mlflow instead and cannot build the container image otherwise, so always pin mlflow explicitly when you call mlflow.pyfunc.log_model().
# DBR ML ships with mlflow-skinny by default, so specify mlflow explicitly
# to ensure Model Serving compatibility.
mlflow.pyfunc.log_model(
name="model",
python_model=your_model,
pip_requirements=["mlflow==3.8.1"], # use mlflow, not mlflow-skinny
registered_model_name="catalog.schema.model_name",
)Checkpoint 4 of 5· Check yourself
A pyfunc model is slow because it loads its weights from disk on every call. Where should the weights be loaded?
Anything that only needs loading once belongs in load_context. That keeps predict, which runs on every request, as light as possible.
“load_context - anything that needs to be loaded just one time for the model to operate should be defined in this function.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
A team has written a Python function in Unity Catalog that looks up current inventory levels by SKU. They want a LangChain tool-calling agent chain to invoke this function as a tool during a conversation. Which approach correctly exposes the Unity Catalog function to the agent?
Correct answer: A — Wrap the function with `UCFunctionToolkit`, converting the registered Unity Catalog function into a LangChain-compatible tool.
- A. Correct: `UCFunctionToolkit` is the documented way to wrap a Unity Catalog function so a LangChain tool-calling agent can discover it and invoke it as a tool when needed.
- B. Incorrect: inlining the function's source and calling it directly bypasses tool-calling entirely; the LLM never gets the chance to decide whether the lookup is needed for a given request.
- C. Incorrect: querying a SQL view manually inside the prompt template runs the lookup on every request regardless of need, which defeats the on-demand invocation that tool-calling agents rely on.
- D. Incorrect: storing output in a Delta table and referencing it as a logging input example only documents the model's schema, it does not give the agent a callable tool during inference.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.In a pyfunc model, load_context runs on every request alongside predict.Why is that wrong?
load_context is for things loaded once. predict contains the logic that runs on each request.
Covered in Wrapping custom logic as a pyfunc model
2.VectorSearchRetrieverTool needs only an index_name, whatever kind of index it points to.Why is that wrong?
Direct-access indexes and Delta Sync indexes with self-managed embeddings also need a custom embedding model and a text column.
Covered in Adding the retrieval step
Practise it for real
Prototype the retrieval component of a chain against an existing index, then pipe a prompt into a Databricks-served LLM so the chain retrieves on every request
1.Run %pip install --upgrade databricks-langchain in a notebook
Why: This package contains Databricks AI Bridge, including VectorSearchRetrieverTool and ChatDatabricks
You should see: The package installs and both classes can be imported
2.Create VectorSearchRetrieverTool with your index_name (catalog.schema.index), a tool_name and a descriptive tool_description
Why: The retriever wraps an existing index. A self-managed embeddings index also needs text_column and embedding
You should see: The tool object is created without errors
3.Call vs_tool.invoke("<a question your documents answer>")
Why: This tests retrieval on its own before any LLM is involved
You should see: Relevant chunks come back from the index
4.Create ChatDatabricks(endpoint="databricks-claude-sonnet-4-5"), a PromptTemplate with placeholders for the question and the retrieved context, and compose prompt_template | llm | StrOutputParser()
Why: This is the pipe-operator chain from the first section, now with the LLM served by Databricks. Your code, not the LLM, decides that retrieval happens first
You should see: The chain object is created without errors
5.Pass the result of vs_tool.invoke and your question into chain.invoke as the template's placeholder values
Why: Retrieve, augment, generate in a fixed order is a deterministic RAG chain. Calling bind_tools instead would let the LLM decide, which is the agent pattern
You should see: A plain string answer that draws on the retrieved documentation
Stuck? Get a nudge
If invoke returns nothing useful, check whether the index computes its own embeddings or needs the embedding argument.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“chain = prompt_template | llm | StrOutputParser()”
↩︎ Composing a chain with LangChain“For production environments, use AI Gateway or Databricks secrets instead of hardcoded values for secure API key management.”
↩︎ Composing a chain with LangChain - 2.
“use Databricks AI Bridge packages like databricks-langchain and databricks-openai”
↩︎ Adding the retrieval step“you must configure the VectorSearchRetrieverTool and specify a custom embedding model and text column.”
↩︎ Exam trap 2 - 3.
“The LLM does not make decisions about which tools to call or in what order.”
↩︎ Adding the retrieval step“Augment a prompt by combining the user request with the retrieved context.”
↩︎ Adding the retrieval step - 4.https://docs.databricks.com/aws/en/machine-learning/model-serving/deploy-custom-python-codeOfficial docs
“MLflow's Python function, pyfunc, provides flexibility to deploy any piece of Python code or any Python model.”
↩︎ Wrapping custom logic as a pyfunc model“Your application requires the model's raw outputs to be post-processed for consumption.”
↩︎ Wrapping custom logic as a pyfunc model“use mlflow.models.predict to validate models before deployment”
↩︎ Wrapping custom logic as a pyfunc model“you can register it to Unity Catalog or Workspace Registry and serve your model to a Model Serving endpoint”
↩︎ Wrapping custom logic as a pyfunc model“Model Serving requires mlflow (not mlflow-skinny) in conda.yaml and cannot build the container image otherwise.”
↩︎ Wrapping custom logic as a pyfunc model“predict - this function houses all the logic that is run every time an input request is made.”
↩︎ Exam trap 1“load_context - anything that needs to be loaded just one time for the model to operate should be defined in this function.”
↩︎ Checkpoint - 5.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-inference-chain-ragOfficial docs
“The series, or chain of steps that are invoked at inference time is commonly referred to as the RAG chain.”
↩︎ Wrapping custom logic as a pyfunc model