CertSafari
    Databricks Certified Generative AI Engineer Associate· Lessons

    Domain 4 · Lesson 31/56

    RAG Application Building Blocks: Flavor, Embeddings, Retriever, Signature

    Choose the basic elements needed to create a RAG application: model flavor, embedding model, retriever, dependencies, input examples, model signature

    16 min read
    1.79% of exam
    10 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Pick an MLflow model flavor (built-in such as LangChain, or custom pyfunc) for a RAG chain
    • Choose an embedding model and keep query-time and index-time embeddings consistent
    • Configure a retriever that returns documents in the MLflow retriever schema
    • Declare a RAG model's dependencies so Model Serving can build its container
    • Attach an input example and a model signature, and explain why Unity Catalog needs the signature

    Key concept

    MLflow Model flavor — A RAG chain gets deployed as one packaged MLflow Model. The flavor you log it with sets how it is saved, which dependencies get captured, and how serving and inference tools load it. The embedding model, retriever, requirements, input example and signature all go into that one package.

    1.The elements a RAG chain needs

    Each element in this objective comes from a step in the RAG chain. Databricks describes the simplest RAG agent in three steps. Retrieval queries an outside knowledge base. Augmentation combines what was retrieved with the user's request in a prompt template. Generation sends that prompt to an LLM. Retrieval in turn needs two things: an embedding model to turn the question into a vector, and a retriever to look up similar chunks. The chain can also have optional pre-processing of the user query and post-processing of the LLM's answer.

    To deploy the chain you package all of it as an MLflow Model. The package needs a flavor, a list of dependencies, an input example and a signature. The table below shows which step each element serves.

    Each RAG element and the part of the chain it serves
    ElementWhat it decidesChain stage
    Embedding modelHow the query and the document chunks become vectorsRetrieval
    RetrieverWhich chunks come back, how many, and in what shapeRetrieval
    Model flavorHow the whole chain is logged and loaded (for example LangChain or custom pyfunc)Packaging the full chain
    DependenciesWhich libraries the serving container must installPackaging / deployment
    Input example and signatureWhat a valid request and response look likePackaging / deployment

    Checkpoint 1 of 8· Put it in order

    Put the RAG chain steps in the order they run for a single user request

    1. 1.Optionally post-process the response (business logic, citations)
    2. 2.Optionally preprocess the user query into a retrieval query
    3. 3.Augment the prompt template with the retrieved context
    4. 4.Generate a response with the LLM
    5. 5.Embed the retrieval query and retrieve the most similar chunks

    Sources12

    2.Choosing the model flavor: built-in (LangChain) or custom pyfunc

    For a framework that has a built-in flavor, log the chain through that flavor's module. MLflow then handles the framework's specifics for you, including the main point of this objective: it reliably records the framework's own dependencies. The example below logs a LangChain chain. It also passes an input example and some parameters, which come up again later in this lesson.

    Logging a LangChain chain with its built-in flavor, including an input examplepython
    model_info = mlflow.langchain.log_model(
      lc_model=chain,
      name="basic_chain",
      params={
        "temperature": 0.1,
        "max_tokens": 2000,
        "prompt_template": str(prompt)
      },
      model_type="agent",
      input_example={"messages": "What is MLflow?"},
    )

    Use the generic python-function (pyfunc) flavor when no built-in flavor fits, or when you need your own pre-processing and post-processing around the chain. Subclass mlflow.pyfunc.PythonModel. Its load_context method sets up the clients and resources, and its predict method runs the logic. Writing a whole custom flavor is also possible, but the docs say it takes more work. Every Python MLflow model, whatever its flavor, can be loaded as a generic Python function with mlflow.pyfunc.load_model().

    Checkpoint 2 of 8· Check yourself

    Your RAG chain needs custom pre-processing of the user query and post-processing that adds citations. No built-in flavor covers that logic. What is the recommended packaging approach?

    Checkpoint 3 of 8· Exam question

    A team is building a RAG chain in Databricks entirely out of LangChain-compatible components: a retriever, a prompt template, and an LLM composed with LangChain's chain APIs. They want to log the chain with MLflow so it is natively serialized, automatically traced, and reloadable with a matching `load_model` call. Which model flavor should they use?

    Sources345

    3.Choosing the embedding model

    No. Retrieval compares the query vector with the chunk vectors, and that comparison only works if both came from the same embedding model. The cookbook says the retrieval query is translated into an embedding using the same embedding model that was used to embed the document chunks during data preparation. So the embedding model is chosen once, for both sides of the comparison.

    When you create a Databricks AI Search index (formerly Vector Search), you also decide who computes the embeddings. That decision limits what you can change later.

    AI Search options for providing embeddings
    OptionSource dataWho computes embeddingsLater change
    Databricks-managed embeddings (Delta Sync)Text in a Delta Lake, streaming, or Iceberg v3+ tableDatabricks, using a model that you specifyIndex stays synced with the source
    Self-managed embeddings (Delta Sync)Table with pre-calculated embeddingsYouCannot be converted to managed; create a new index and recompute
    Direct Vector AccessEmbeddings you write via APIYouYou must manually update the index using the REST API

    If none of the hosted embedding endpoints suits you, you can bring your own model, packaged with the same MLflow mechanics as the chain. Log it as a custom Python model. Its input is a single string column, and its output is a tensor of embeddings returned as a NumPy array.

    Checkpoint 4 of 8· Check yourself

    A team built an AI Search index with self-managed embeddings. They now want Databricks to compute embeddings for them. What must they do?

    Sources267

    4.Configuring the retriever

    The retriever is the component the chain calls to fetch supporting chunks. For Databricks AI Search indexes, the AI Bridge packages databricks-langchain and databricks-openai provide VectorSearchRetrieverTool, plus helpers such as from_vector_search and from_uc_function that build retrievers from existing Databricks resources. On the tool you set how many results to return, which columns to return, filters, and the query type: ANN or HYBRID. Hybrid search helps when the source data contains exact keywords such as SKUs or identifiers.

    The embedding model comes up again here. With a managed-embedding index, Databricks embeds the query for you. With a direct-access index, or a Delta Sync index using self-managed embeddings, the retriever cannot know which model to use, so you have to pass it the embedding model and the text column.

    A retriever for a self-managed-embedding index, with its embedding model and text column passed explicitlypython
    embedding_model = DatabricksEmbeddings(
        endpoint="databricks-bge-large-en",
    )
    
    vs_tool = VectorSearchRetrieverTool(
      index_name="catalog.schema.index_name", # Index name in the format 'catalog.schema.index'
      num_results=5, # Max number of documents to return
      columns=["primary_key", "text_column"], # List of columns to include in the search
      filters={"text_column LIKE": "Databricks"}, # Filters to apply to the query
      query_type="ANN", # Query type ("ANN" or "HYBRID").
      tool_name="name of the tool", # Used by the LLM to understand the purpose of the tool
      tool_description="Purpose of the tool", # Used by the LLM to understand the purpose of the tool
      text_column="text_column", # Specify text column for embeddings. Required for direct-access index or delta-sync index with self-managed embeddings.
      embedding=embedding_model # The embedding model. Required for direct-access index or delta-sync index with self-managed embeddings.
    )

    If you write your own retriever, for example in a pyfunc-flavored agent that calls a vector store hosted outside Databricks, what it returns matters too. To fit the MLflow retriever schema, it should return a list of MLflow Document objects. Extra attributes such as doc_uri or a similarity score go in each Document's metadata field.

    Checkpoint 5 of 8· Check yourself

    You write a custom retriever function for a pyfunc-flavored agent. What should it return to conform to the MLflow retriever schema, and where do extra attributes such as a similarity score go?

    Sources86

    5.Declaring dependencies

    During deployment, Model Serving builds a production container from what the MLflow model declares. The base image may include some system-level packages, but your application-level dependencies must be declared explicitly in the model. Logging a model automatically writes requirements.txt and conda.yaml. With mlflow.pyfunc.log_model, MLflow infers the requirements using mlflow.models.infer_pip_requirements. Inference doesn't always catch everything, so MLflow gives you several ways to fill the gaps.

    Ways to control a model's dependencies when logging it
    MechanismUse it forNote
    Automatic inference (infer_pip_requirements)Default for pyfunc and built-in flavorsWritten to requirements.txt as a model artifact
    extra_pip_requirementsAdding libraries that inference missedKeeps the inferred set
    pip_requirements / conda_envReplacing the whole requirement setGenerally discouraged; overrides what MLflow picks up
    code_paths (code_path in MLflow 2.x)Local .py files or custom wheelsAdded to the Python path at load time
    artifactsA .txt or .json listing non-Python dependenciesMLflow does not pick these up automatically
    MLFLOW_LOCK_MODEL_DEPENDENCIESAlso capturing transitive dependenciesMLflow 3

    Pin exact versions, for example nltk==<version> rather than just nltk. One trap is specific to Databricks. Databricks Runtime ML includes mlflow-skinny rather than the full mlflow package. If you log a pyfunc model there without pip_requirements, the model's conda.yaml records mlflow-skinny, and Model Serving cannot build the container image because it needs mlflow. The fix is to pin mlflow==<version> in pip_requirements. Databricks also recommends running inference on the same runtime version you used to create the model.

    Checkpoint 6 of 8· Fill the gap

    Your chain imports a helper module, utils.py, that cannot be installed with pip. Which parameter (MLflow 3) packages it with the model so it is on the Python path when the model loads?

    mlflow.pyfunc.log_model(
       name=name,
        ? =[filename.py],
       data_path=data_path,
       conda_env=conda_env,
    )

    Checkpoint 7 of 8· Exam question

    A developer is creating a Delta Sync index in Mosaic AI Vector Search (Databricks AI Search) over a Delta table of chunked documents. They want Databricks to automatically compute and refresh embeddings from a text column whenever the source table changes, without operating a separate embedding pipeline. Which embedding configuration satisfies this?

    Sources94

    6.Input examples and model signatures

    The last two elements describe the model's interface. An input example is a sample request saved with the model, such as the {"messages": "What is MLflow?"} passed to the LangChain log_model call earlier. A signature formally declares the input and output schema. One way to get a signature is to infer it from sample data and the model's output, then pass it to log_model, as the custom models overview shows.

    You can also build the signature by hand. The custom embedding model from earlier does this, and the same call brings together the elements this lesson has covered. The input schema is one string column (ColSpec) and the output is a float32 tensor (TensorSpec). The model uses the pyfunc flavor, its requirements are pinned and include mlflow==<version> as discussed above, and it carries both a signature and an input example.

    A pyfunc embedding model logged with an explicit signature, pinned requirements and an input examplepython
    with mlflow.start_run() as run:
        input_schema = Schema([ColSpec(DataType.string)])
        output_schema = Schema([TensorSpec(np.dtype(np.float32), [-1, 1024])])
        signature = ModelSignature(inputs=input_schema, outputs=output_schema)
        mlflow.pyfunc.log_model(
            artifact_path="model",
            python_model=CustomEmbeddingModel(),
            pip_requirements=["mlflow==2.21.3", "openai==1.69.0"],
            signature=signature,
            input_example=input_example
        )

    The signature is not optional if the model is going to Unity Catalog. The documentation states plainly that Models in Unity Catalog require a signature, and model versions registered without one have limitations. Registering and serving the model are covered in neighbouring lessons. For this objective, what matters is to log the signature with the chain from the start.

    Checkpoint 8 of 8· Check yourself

    You are logging a custom embedding model for AI Search. Which signature does the documentation call for?

    Sources107

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.You can embed user queries with any embedding model, even a different one from the model that built the index.Why is that wrong?

      Similarity is only meaningful when the query and the chunks were embedded by the same model, so the query must use the index's embedding model.

      Covered in Choosing the embedding model

    2. 2.MLflow's automatic inference captures every dependency, so logging a pyfunc model on Databricks Runtime ML with default settings is always servable.Why is that wrong?

      On ML runtimes MLflow records mlflow-skinny, and Model Serving cannot build the image unless you pin mlflow in pip_requirements. Non-Python dependencies are never picked up automatically either.

      Covered in Declaring dependencies

    3. 3.A model signature is optional metadata. An input example alone is enough to register the chain in Unity Catalog.Why is that wrong?

      Unity Catalog requires a signature, and versions registered without one have limitations.

      Covered in Input examples and model signatures

    4. 4.VectorSearchRetrieverTool always embeds the query for you, whatever kind of index it points at.Why is that wrong?

      For direct-access indexes and Delta Sync indexes with self-managed embeddings, you must supply the embedding model and the text column yourself.

      Covered in Configuring the retriever

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Retrieval: The user's request is used to query an outside knowledge base such as a vector store, keyword search, or SQL database.”
      ↩︎ The elements a RAG chain needs
    2. 2.
      “The LLM's response may be processed further to apply additional business logic, add citations, or otherwise refine the generated text”
      ↩︎ The elements a RAG chain needs
      “the retrieval query is translated into an embedding using the same embedding model that was used to embed the document chunks during data preparation”
      ↩︎ Choosing the embedding model
      “the retrieval query is translated into an embedding using the same embedding model that was used to embed the document chunks during data preparation”
      ↩︎ Exam trap 1
    3. 4.
      “MLflow has native support for some Python ML libraries, where MLflow can reliably log dependencies for models that use these libraries.”
      ↩︎ Choosing the model flavor: built-in (LangChain) or custom pyfunc
      “MLflow infers the dependencies using mlflow.models.infer_pip_requirements, and logs them to a requirements.txt file as a model artifact.”
      ↩︎ Declaring dependencies
      “doing so is generally discouraged because this overrides the dependencies which MLflow picks up automatically”
      ↩︎ Declaring dependencies
      “MLflow does not automatically pick up non-Python dependencies, such as Java packages, R packages, and native packages (such as Linux packages).”
      ↩︎ Declaring dependencies
      “Model Serving cannot build the container image because it requires mlflow”
      ↩︎ Declaring dependencies
      “Databricks recommends running inference on the same runtime version that you used during training”
      ↩︎ Declaring dependencies
      “Specify mlflow==<version> in pip_requirements when you log the model.”
      ↩︎ Exam trap 2
      “Write a custom Python model. Doing so allows you to subclass mlflow.pyfunc.PythonModel to customize initialization and prediction.”
      ↩︎ Checkpoint
    4. 5.
      “For Python MLflow models, an additional option is to use mlflow.pyfunc.load_model() to load the model as a generic Python function.”
      ↩︎ Choosing the model flavor: built-in (LangChain) or custom pyfunc
      “The format defines a convention that lets you save a model in different flavors (python-function, pytorch, sklearn, and so on)”
      ↩︎ Key concept
    5. 6.
      “Databricks calculates the embeddings, using a model that you specify, and optionally saves the embeddings to a table in Unity Catalog.”
      ↩︎ Choosing the embedding model
      “This method is particularly useful in RAG applications where source data has unique keywords such as SKUs or identifiers”
      ↩︎ Configuring the retriever
      “It is not possible to convert a self-managed embedding index to a Databricks-managed index.”
      ↩︎ Checkpoint
    6. 7.
      “To use your own embedding model with AI Search, log it as a custom Python model with MLflow.”
      ↩︎ Choosing the embedding model
      “The input schema should be a single ColSpec string, and the output schema should be TensorSpec as your signature.”
      ↩︎ Input examples and model signatures
    7. 8.
      “These packages include helper functions like from_vector_search and from_uc_function to create retrievers from existing Databricks resources.”
      ↩︎ Configuring the retriever
      “the retriever function should return a List[Document] object and use the metadata field in the Document class”
      ↩︎ Configuring the retriever
      “you must configure the VectorSearchRetrieverTool and specify a custom embedding model and text column”
      ↩︎ Exam trap 4
    8. 9.
      “application-level dependencies must be explicitly specified in your MLflow model.”
      ↩︎ Declaring dependencies
    9. 10.
      “Models in Unity Catalog require a signature.”
      ↩︎ Input examples and model signatures
      “Models in Unity Catalog require a signature.”
      ↩︎ Exam trap 3

    Spotted a mistake, or was something unclear? Tell us.