What you will be able to do
- Pick an MLflow model flavor (built-in such as LangChain, or custom pyfunc) for a RAG chain
- Choose an embedding model and keep query-time and index-time embeddings consistent
- Configure a retriever that returns documents in the MLflow retriever schema
- Declare a RAG model's dependencies so Model Serving can build its container
- Attach an input example and a model signature, and explain why Unity Catalog needs the signature
Key concept
MLflow Model flavor — A RAG chain gets deployed as one packaged MLflow Model. The flavor you log it with sets how it is saved, which dependencies get captured, and how serving and inference tools load it. The embedding model, retriever, requirements, input example and signature all go into that one package.
1.The elements a RAG chain needs
Each element in this objective comes from a step in the RAG chain. Databricks describes the simplest RAG agent in three steps. Retrieval queries an outside knowledge base. Augmentation combines what was retrieved with the user's request in a prompt template. Generation sends that prompt to an LLM. Retrieval in turn needs two things: an embedding model to turn the question into a vector, and a retriever to look up similar chunks. The chain can also have optional pre-processing of the user query and post-processing of the LLM's answer.
To deploy the chain you package all of it as an MLflow Model. The package needs a flavor, a list of dependencies, an input example and a signature. The table below shows which step each element serves.
| Element | What it decides | Chain stage |
|---|---|---|
| Embedding model | How the query and the document chunks become vectors | Retrieval |
| Retriever | Which chunks come back, how many, and in what shape | Retrieval |
| Model flavor | How the whole chain is logged and loaded (for example LangChain or custom pyfunc) | Packaging the full chain |
| Dependencies | Which libraries the serving container must install | Packaging / deployment |
| Input example and signature | What a valid request and response look like | Packaging / deployment |
Checkpoint 1 of 8· Put it in order
Put the RAG chain steps in the order they run for a single user request
- 1.Optionally post-process the response (business logic, citations)
- 2.Optionally preprocess the user query into a retrieval query
- 3.Augment the prompt template with the retrieved context
- 4.Generate a response with the LLM
- 5.Embed the retrieval query and retrieve the most similar chunks
The RAG chain runs preprocessing, retrieval, prompt augmentation, LLM generation and then optional post-processing, in that order.
“The LLM's response may be processed further to apply additional business logic, add citations, or otherwise refine the generated text”Source: docs.databricks.com
2.Choosing the model flavor: built-in (LangChain) or custom pyfunc
For a framework that has a built-in flavor, log the chain through that flavor's module. MLflow then handles the framework's specifics for you, including the main point of this objective: it reliably records the framework's own dependencies. The example below logs a LangChain chain. It also passes an input example and some parameters, which come up again later in this lesson.
model_info = mlflow.langchain.log_model(
lc_model=chain,
name="basic_chain",
params={
"temperature": 0.1,
"max_tokens": 2000,
"prompt_template": str(prompt)
},
model_type="agent",
input_example={"messages": "What is MLflow?"},
)Use the generic python-function (pyfunc) flavor when no built-in flavor fits, or when you need your own pre-processing and post-processing around the chain. Subclass mlflow.pyfunc.PythonModel. Its load_context method sets up the clients and resources, and its predict method runs the logic. Writing a whole custom flavor is also possible, but the docs say it takes more work. Every Python MLflow model, whatever its flavor, can be loaded as a generic Python function with mlflow.pyfunc.load_model().
Checkpoint 2 of 8· Check yourself
Your RAG chain needs custom pre-processing of the user query and post-processing that adds citations. No built-in flavor covers that logic. What is the recommended packaging approach?
A custom Python model lets you customize initialization and prediction, and it works well for Python-only customization. A custom flavor takes more work to implement.
“Write a custom Python model. Doing so allows you to subclass mlflow.pyfunc.PythonModel to customize initialization and prediction.”Source: docs.databricks.com
Checkpoint 3 of 8· Exam question
A team is building a RAG chain in Databricks entirely out of LangChain-compatible components: a retriever, a prompt template, and an LLM composed with LangChain's chain APIs. They want to log the chain with MLflow so it is natively serialized, automatically traced, and reloadable with a matching `load_model` call. Which model flavor should they use?
Correct answer: A — The `mlflow.langchain` flavor, since it natively serializes LangChain objects and integrates with MLflow tracing for chain-based components like retrievers and prompt templates.
- A. The LangChain flavor is built specifically to serialize LangChain chain objects such as retrievers, prompt templates, and LLM wrappers, and it plugs directly into MLflow's tracing and autologging for chain components. This matches a chain composed entirely of LangChain-native pieces.
- B. The generic PyFunc flavor is a valid way to log custom code, but for a chain already built from LangChain components it discards the native chain serialization and forces manual reimplementation of the prediction logic, which is unnecessary extra work here.
- C. This flavor targets scikit-learn estimator objects trained with `fit`, which has no relationship to a retrieval-augmented LangChain chain and cannot serialize LangChain's chain graph.
- D. This flavor is designed for Hugging Face pipeline objects and would only capture a single text-generation step, losing the retriever and prompt-template composition that makes up the rest of the chain.
3.Choosing the embedding model
No. Retrieval compares the query vector with the chunk vectors, and that comparison only works if both came from the same embedding model. The cookbook says the retrieval query is translated into an embedding using the same embedding model that was used to embed the document chunks during data preparation. So the embedding model is chosen once, for both sides of the comparison.
When you create a Databricks AI Search index (formerly Vector Search), you also decide who computes the embeddings. That decision limits what you can change later.
| Option | Source data | Who computes embeddings | Later change |
|---|---|---|---|
| Databricks-managed embeddings (Delta Sync) | Text in a Delta Lake, streaming, or Iceberg v3+ table | Databricks, using a model that you specify | Index stays synced with the source |
| Self-managed embeddings (Delta Sync) | Table with pre-calculated embeddings | You | Cannot be converted to managed; create a new index and recompute |
| Direct Vector Access | Embeddings you write via API | You | You must manually update the index using the REST API |
If none of the hosted embedding endpoints suits you, you can bring your own model, packaged with the same MLflow mechanics as the chain. Log it as a custom Python model. Its input is a single string column, and its output is a tensor of embeddings returned as a NumPy array.
Checkpoint 4 of 8· Check yourself
A team built an AI Search index with self-managed embeddings. They now want Databricks to compute embeddings for them. What must they do?
A self-managed embedding index cannot be converted to a managed one. Switching means building a new index.
“It is not possible to convert a self-managed embedding index to a Databricks-managed index.”Source: docs.databricks.com
4.Configuring the retriever
The retriever is the component the chain calls to fetch supporting chunks. For Databricks AI Search indexes, the AI Bridge packages databricks-langchain and databricks-openai provide VectorSearchRetrieverTool, plus helpers such as from_vector_search and from_uc_function that build retrievers from existing Databricks resources. On the tool you set how many results to return, which columns to return, filters, and the query type: ANN or HYBRID. Hybrid search helps when the source data contains exact keywords such as SKUs or identifiers.
The embedding model comes up again here. With a managed-embedding index, Databricks embeds the query for you. With a direct-access index, or a Delta Sync index using self-managed embeddings, the retriever cannot know which model to use, so you have to pass it the embedding model and the text column.
embedding_model = DatabricksEmbeddings(
endpoint="databricks-bge-large-en",
)
vs_tool = VectorSearchRetrieverTool(
index_name="catalog.schema.index_name", # Index name in the format 'catalog.schema.index'
num_results=5, # Max number of documents to return
columns=["primary_key", "text_column"], # List of columns to include in the search
filters={"text_column LIKE": "Databricks"}, # Filters to apply to the query
query_type="ANN", # Query type ("ANN" or "HYBRID").
tool_name="name of the tool", # Used by the LLM to understand the purpose of the tool
tool_description="Purpose of the tool", # Used by the LLM to understand the purpose of the tool
text_column="text_column", # Specify text column for embeddings. Required for direct-access index or delta-sync index with self-managed embeddings.
embedding=embedding_model # The embedding model. Required for direct-access index or delta-sync index with self-managed embeddings.
)If you write your own retriever, for example in a pyfunc-flavored agent that calls a vector store hosted outside Databricks, what it returns matters too. To fit the MLflow retriever schema, it should return a list of MLflow Document objects. Extra attributes such as doc_uri or a similarity score go in each Document's metadata field.
Checkpoint 5 of 8· Check yourself
You write a custom retriever function for a pyfunc-flavored agent. What should it return to conform to the MLflow retriever schema, and where do extra attributes such as a similarity score go?
The MLflow retriever schema expects a list of Document objects, with additional attributes such as doc_uri or similarity_score carried in each Document's metadata field.
“the retriever function should return a List[Document] object and use the metadata field in the Document class”Source: docs.databricks.com
Databricks already computes the embeddings with the model you chose when you created the index, so it can embed the query text itself. The docs require these two arguments only for direct-access indexes and Delta Sync indexes with self-managed embeddings.
5.Declaring dependencies
During deployment, Model Serving builds a production container from what the MLflow model declares. The base image may include some system-level packages, but your application-level dependencies must be declared explicitly in the model. Logging a model automatically writes requirements.txt and conda.yaml. With mlflow.pyfunc.log_model, MLflow infers the requirements using mlflow.models.infer_pip_requirements. Inference doesn't always catch everything, so MLflow gives you several ways to fill the gaps.
| Mechanism | Use it for | Note |
|---|---|---|
| Automatic inference (infer_pip_requirements) | Default for pyfunc and built-in flavors | Written to requirements.txt as a model artifact |
| extra_pip_requirements | Adding libraries that inference missed | Keeps the inferred set |
| pip_requirements / conda_env | Replacing the whole requirement set | Generally discouraged; overrides what MLflow picks up |
| code_paths (code_path in MLflow 2.x) | Local .py files or custom wheels | Added to the Python path at load time |
| artifacts | A .txt or .json listing non-Python dependencies | MLflow does not pick these up automatically |
| MLFLOW_LOCK_MODEL_DEPENDENCIES | Also capturing transitive dependencies | MLflow 3 |
Pin exact versions, for example nltk==<version> rather than just nltk. One trap is specific to Databricks. Databricks Runtime ML includes mlflow-skinny rather than the full mlflow package. If you log a pyfunc model there without pip_requirements, the model's conda.yaml records mlflow-skinny, and Model Serving cannot build the container image because it needs mlflow. The fix is to pin mlflow==<version> in pip_requirements. Databricks also recommends running inference on the same runtime version you used to create the model.
Checkpoint 6 of 8· Fill the gap
Your chain imports a helper module, utils.py, that cannot be installed with pip. Which parameter (MLflow 3) packages it with the model so it is on the Python path when the model loads?
mlflow.pyfunc.log_model(
name=name,
? =[filename.py],
data_path=data_path,
conda_env=conda_env,
)code_paths stores the files in the model's code directory, and MLflow adds them to the Python path when the model loads. The pip parameters only list installable packages.
Source: docs.databricks.comCheckpoint 7 of 8· Exam question
A developer is creating a Delta Sync index in Mosaic AI Vector Search (Databricks AI Search) over a Delta table of chunked documents. They want Databricks to automatically compute and refresh embeddings from a text column whenever the source table changes, without operating a separate embedding pipeline. Which embedding configuration satisfies this?
Correct answer: A — Specify `embedding_source_column` pointing to the text column and select a Databricks-hosted embedding model, letting the index compute embeddings automatically during each sync.
- A. Pointing the index at the text column with `embedding_source_column` and choosing a Databricks-hosted embedding model lets Databricks compute embeddings automatically on every sync, which is exactly the hands-off refresh behavior the developer wants.
- B. This describes the self-managed embeddings path, where the caller supplies precomputed vectors rather than letting Databricks compute them from the text column, so it does not remove the need for an external embedding pipeline.
- C. Manually invoking a model before every table update reintroduces the external pipeline the developer wants to avoid, since embeddings are no longer computed automatically as part of the sync.
- D. A Direct Vector Access index expects vectors and metadata to be pushed manually through the API rather than syncing and embedding automatically from a Delta table, which is the opposite of the automatic behavior requested.
6.Input examples and model signatures
The last two elements describe the model's interface. An input example is a sample request saved with the model, such as the {"messages": "What is MLflow?"} passed to the LangChain log_model call earlier. A signature formally declares the input and output schema. One way to get a signature is to infer it from sample data and the model's output, then pass it to log_model, as the custom models overview shows.
You can also build the signature by hand. The custom embedding model from earlier does this, and the same call brings together the elements this lesson has covered. The input schema is one string column (ColSpec) and the output is a float32 tensor (TensorSpec). The model uses the pyfunc flavor, its requirements are pinned and include mlflow==<version> as discussed above, and it carries both a signature and an input example.
with mlflow.start_run() as run:
input_schema = Schema([ColSpec(DataType.string)])
output_schema = Schema([TensorSpec(np.dtype(np.float32), [-1, 1024])])
signature = ModelSignature(inputs=input_schema, outputs=output_schema)
mlflow.pyfunc.log_model(
artifact_path="model",
python_model=CustomEmbeddingModel(),
pip_requirements=["mlflow==2.21.3", "openai==1.69.0"],
signature=signature,
input_example=input_example
)The signature is not optional if the model is going to Unity Catalog. The documentation states plainly that Models in Unity Catalog require a signature, and model versions registered without one have limitations. Registering and serving the model are covered in neighbouring lessons. For this objective, what matters is to log the signature with the chain from the start.
Checkpoint 8 of 8· Check yourself
You are logging a custom embedding model for AI Search. Which signature does the documentation call for?
The embedding model takes text in and returns vectors, so the input is a single string column and the output is a tensor, returned as a NumPy array.
“The input schema should be a single ColSpec string, and the output schema should be TensorSpec as your signature.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.You can embed user queries with any embedding model, even a different one from the model that built the index.Why is that wrong?
Similarity is only meaningful when the query and the chunks were embedded by the same model, so the query must use the index's embedding model.
Covered in Choosing the embedding model
2.MLflow's automatic inference captures every dependency, so logging a pyfunc model on Databricks Runtime ML with default settings is always servable.Why is that wrong?
On ML runtimes MLflow records mlflow-skinny, and Model Serving cannot build the image unless you pin mlflow in pip_requirements. Non-Python dependencies are never picked up automatically either.
Covered in Declaring dependencies
3.A model signature is optional metadata. An input example alone is enough to register the chain in Unity Catalog.Why is that wrong?
Unity Catalog requires a signature, and versions registered without one have limitations.
Covered in Input examples and model signatures
4.VectorSearchRetrieverTool always embeds the query for you, whatever kind of index it points at.Why is that wrong?
For direct-access indexes and Delta Sync indexes with self-managed embeddings, you must supply the embedding model and the text column yourself.
Covered in Configuring the retriever
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Retrieval: The user's request is used to query an outside knowledge base such as a vector store, keyword search, or SQL database.”
↩︎ The elements a RAG chain needs - 2.https://docs.databricks.com/aws/en/agents/tutorials/ai-cookbook/fundamentals-inference-chain-ragOfficial docs
“The LLM's response may be processed further to apply additional business logic, add citations, or otherwise refine the generated text”
↩︎ The elements a RAG chain needs“the retrieval query is translated into an embedding using the same embedding model that was used to embed the document chunks during data preparation”
↩︎ Choosing the embedding model“the retrieval query is translated into an embedding using the same embedding model that was used to embed the document chunks during data preparation”
↩︎ Exam trap 1 - 3.
“Use the log_model() method for your flavor of agent.”
↩︎ Choosing the model flavor: built-in (LangChain) or custom pyfunc - 4.
“MLflow has native support for some Python ML libraries, where MLflow can reliably log dependencies for models that use these libraries.”
↩︎ Choosing the model flavor: built-in (LangChain) or custom pyfunc“MLflow infers the dependencies using mlflow.models.infer_pip_requirements, and logs them to a requirements.txt file as a model artifact.”
↩︎ Declaring dependencies“doing so is generally discouraged because this overrides the dependencies which MLflow picks up automatically”
↩︎ Declaring dependencies“MLflow does not automatically pick up non-Python dependencies, such as Java packages, R packages, and native packages (such as Linux packages).”
↩︎ Declaring dependencies“Model Serving cannot build the container image because it requires mlflow”
↩︎ Declaring dependencies“Databricks recommends running inference on the same runtime version that you used during training”
↩︎ Declaring dependencies“Specify mlflow==<version> in pip_requirements when you log the model.”
↩︎ Exam trap 2“Write a custom Python model. Doing so allows you to subclass mlflow.pyfunc.PythonModel to customize initialization and prediction.”
↩︎ Checkpoint - 5.https://docs.databricks.com/aws/en/mlflow/modelsOfficial docs
“For Python MLflow models, an additional option is to use mlflow.pyfunc.load_model() to load the model as a generic Python function.”
↩︎ Choosing the model flavor: built-in (LangChain) or custom pyfunc“The format defines a convention that lets you save a model in different flavors (python-function, pytorch, sklearn, and so on)”
↩︎ Key concept - 6.
“Databricks calculates the embeddings, using a model that you specify, and optionally saves the embeddings to a table in Unity Catalog.”
↩︎ Choosing the embedding model“This method is particularly useful in RAG applications where source data has unique keywords such as SKUs or identifiers”
↩︎ Configuring the retriever“It is not possible to convert a self-managed embedding index to a Databricks-managed index.”
↩︎ Checkpoint - 7.
“To use your own embedding model with AI Search, log it as a custom Python model with MLflow.”
↩︎ Choosing the embedding model“The input schema should be a single ColSpec string, and the output schema should be TensorSpec as your signature.”
↩︎ Input examples and model signatures - 8.
“These packages include helper functions like from_vector_search and from_uc_function to create retrievers from existing Databricks resources.”
↩︎ Configuring the retriever“the retriever function should return a List[Document] object and use the metadata field in the Document class”
↩︎ Configuring the retriever“you must configure the VectorSearchRetrieverTool and specify a custom embedding model and text column”
↩︎ Exam trap 4 - 9.
“application-level dependencies must be explicitly specified in your MLflow model.”
↩︎ Declaring dependencies - 10.https://docs.databricks.com/aws/en/machine-learning/manage-model-lifecycle/migrate-modelsOfficial docs
“Models in Unity Catalog require a signature.”
↩︎ Input examples and model signatures“Models in Unity Catalog require a signature.”
↩︎ Exam trap 3