What you will be able to do
- Package files a pyfunc chain needs with the artifacts parameter and read them through context.artifacts
- Ship shared preprocessing code with code_paths (MLflow 3) or code_path (MLflow 2.x)
- Pin dependencies so Model Serving can build the container, including the mlflow vs mlflow-skinny issue
- Add a signature and input example, load the logged chain with mlflow.pyfunc.load_model, and validate it before serving
1.Packaging the files your chain loads
A pyfunc chain usually needs files when it runs, such as model weights or a tokenizer cache. In notebooks these often sit in Unity Catalog volumes, and some models download pieces from the internet, for example HuggingFace tokenizers. Databricks notes that "Real-time workloads at scale perform best when all required dependencies are statically captured at deployment time." Files in volumes therefore must be packaged, and network artifacts should be packaged whenever possible.
You package them with the artifacts parameter of log_model(). It is a dictionary that maps a name to a path. At load time, each path is available inside the model under context.artifacts[<name>], which is why the load_context in a pyfunc model reads context.artifacts["model-weights"]. The names you pick when logging are the keys your code uses when loading.
Checkpoint 1 of 5· Fill the gap
Which parameter packages these files into the logged model?
mlflow.pyfunc.log_model(
...
? ={'model-weights': "/Volumes/catalog/schema/volume/path/to/file", "tokenizer_cache": "./tokenizer_cache"},
...
)artifacts takes a mapping from a name to a file path. Each file is copied into the model, and load_context reads it back as context.artifacts[name].
Source: docs.databricks.comSources1
2.Shipping shared preprocessing code
Preprocessing logic often lives in your own .py modules, like a custom tokenizer or prompt-formatting helpers, and %pip install can't install those. To include them, you pass the paths when you log. Authors "can log full code references that load into the path," so a model logged with code_path = ["preprocessing_utils/"] can import from preprocessing_utils inside its own methods. MLflow stores these files in a code directory next to the model. When the model loads, "MLflow adds these files or directories to the Python path." Custom Python wheel files can be shipped the same way.
The parameter name depends on the MLflow version, and so does the name of the argument that identifies the model:
| Purpose | MLflow 3 | MLflow 2.x |
|---|---|---|
| Identify the logged model | name=name | artifact_path=artifact_path |
| Ship local code (.py files, directories, wheels) | code_paths=[filename.py] | code_path=[filename.py] |
Checkpoint 2 of 5· Check yourself
You are logging a pyfunc chain with MLflow 3 and need to ship a local preprocessing module. Which parameter of mlflow.pyfunc.log_model do you use?
MLflow 3 uses code_paths, while code_path is the MLflow 2.x name. Both ship local files with the model and add them to the Python path at load time.
“using the code_paths parameter (or code_path in MLflow 2.x)”Source: docs.databricks.com
The preprocessing_utils module was shipped with code_path. MLflow adds it to the Python path when the model loads, so importing it inside load_context makes sure the packaged copy is the one used, in whatever environment the model is loaded.
3.Pinning dependencies so serving can build the image
When you log a model, MLflow automatically writes requirements.txt and conda.yaml files. For mlflow.pyfunc.log_model, MLflow tries to work out the dependencies itself with mlflow.models.infer_pip_requirements. A custom chain imports libraries that MLflow has no flavor for, so check what got captured, and log requirements with exact versions (f"nltk=={nltk.__version__}", not just nltk).
You can control the result in two ways:
- **extra_pip_requirements adds packages that inference missed.
- pip_requirements or conda_env** replaces the whole set. This is "generally discouraged because this overrides the dependencies which MLflow picks up automatically."
One case makes overriding mandatory. Databricks Runtime ML ships mlflow-skinny instead of the full mlflow package. If you log a pyfunc model there without pip_requirements, mlflow-skinny is what ends up in conda.yaml. Model Serving requires mlflow and can't build the container image without it.
# DBR ML ships with mlflow-skinny by default, so specify mlflow explicitly
# to ensure Model Serving compatibility.
mlflow.pyfunc.log_model(
name="model",
python_model=your_model,
pip_requirements=["mlflow==3.8.1"], # use mlflow, not mlflow-skinny
registered_model_name="catalog.schema.model_name",
)Checkpoint 3 of 5· Check yourself
A pyfunc chain logged on Databricks Runtime ML without pip_requirements works in the notebook, but the serving endpoint can't build its container. What is the most likely cause?
Databricks Runtime ML includes mlflow-skinny by default, and that is what gets recorded in conda.yaml. Model Serving can't build the image without the full mlflow package, so pin mlflow==<version> in pip_requirements.
“Model Serving requires mlflow (not mlflow-skinny) in conda.yaml and cannot build the container image otherwise.”Source: docs.databricks.com
Checkpoint 4 of 5· Exam question
A team is building a custom pyfunc chain that calls a proprietary tokenizer stored in a private Python package used only by their team, plus a serialized vocabulary file that the tokenizer depends on at runtime. When calling `mlflow.pyfunc.log_model` for this chain, which combination of parameters correctly makes both available inside the deployed model?
Correct answer: A — Pass the vocabulary file path via the `artifacts` parameter and the private package directory via the `code_paths` parameter.
- A. This is correct: `artifacts` is designed for data files like a serialized vocabulary that get copied alongside the model and accessed via `context.artifacts` in `load_context`, while `code_paths` bundles custom local source code, such as a private package, so it is importable inside the serving environment.
- B. `pip_requirements` specifies installable package dependencies with version constraints, not local file or directory paths, so pointing it at a vocabulary file or a private package directory would not make either available to the model.
- C. This swaps the two mechanisms: a data file like a vocabulary is not Python source code and does not belong in `code_paths`, while a package directory of source code is not a data artifact and would not be handled correctly by `artifacts`.
- D. `input_example` is used to supply a sample input so MLflow can infer and validate a model signature, and it has no role in bundling dependency files or private source packages with the logged model.
4.Signature, input example, and testing the chain before serving
Databricks recommends adding a signature and an input example. The signature isn't optional if you want the chain in Unity Catalog, because "Signatures are necessary for logging models to the Unity Catalog." infer_signature works it out from sample inputs and the model's outputs on them:
from mlflow.models.signature import infer_signature
signature = infer_signature(training_data, model.predict(training_data))
mlflow.sklearn.log_model(model, "model", signature=signature)These examples use the sklearn flavor, but mlflow.pyfunc.log_model takes the same signature and input_example keywords. The input_example is a sample request, for example {"feature1": 0.5, "feature2": 3}.
Next, test the logged chain the same way a consumer would call it. You can load any Python MLflow model with mlflow.pyfunc.load_model() as a generic Python function and call .predict(). That runs load_context and then your pre-process, model, post-process path. Databricks also suggests checking that the model can actually be served before you deploy it, using mlflow.models.predict. Once it passes, you can register it to Unity Catalog or the Workspace Registry and serve it from a Model Serving endpoint. Those steps belong to a separate objective.
model = mlflow.pyfunc.load_model(model_path)
model.predict(model_input)Checkpoint 5 of 5· Match them up
Match each log_model parameter to the problem it solves for a pyfunc chain.
Tap a term, then the definition that fits it.
Each parameter covers a different thing the chain depends on: files, local code, installable packages, and its input and output schema.
“Signatures are necessary for logging models to the Unity Catalog.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.MLflow's automatic dependency capture on Databricks Runtime ML is always good enough for Model Serving.Why is that wrong?
Databricks Runtime ML includes mlflow-skinny, which is what conda.yaml records. Model Serving needs the full mlflow package, so pin it in pip_requirements.
Covered in Pinning dependencies so serving can build the image
2.A served chain can keep reading its weights and tokenizer files straight from a Unity Catalog volume path.Why is that wrong?
Model Serving requires files in volumes to be packaged into the model with the artifacts parameter. load_context then reads them through context.artifacts.
Covered in Packaging the files your chain loads
3.The parameter for shipping local .py modules is code_path in every MLflow version.Why is that wrong?
MLflow 3 uses code_paths. code_path is the MLflow 2.x name.
Covered in Shipping shared preprocessing code
Practise it for real
Log a pyfunc chain with preprocessing and postprocessing so that it loads and predicts the way a serving endpoint would call it.
1.Subclass mlflow.pyfunc.PythonModel. Load the files you need in load_context, and in predict call format_inputs, then the model, then format_outputs.
Why: One-time loading goes in load_context. Per-request pre- and post-processing goes in predict.
You should see: A class whose predict returns post-processed output rather than raw model output.
2.Call mlflow.pyfunc.log_model with python_model set to your class, artifacts pointing at your weights or tokenizer files, and code_paths listing any local helper modules.
Why: Model Serving needs files and local code packaged inside the model artifact.
You should see: A logged model whose artifacts include your files and a code directory.
3.In the same log_model call, set pip_requirements with an exact mlflow==<version> pin, and add a signature and input_example.
Why: Databricks Runtime ML would otherwise record mlflow-skinny, and Unity Catalog requires a signature.
You should see: A conda.yaml that lists mlflow rather than mlflow-skinny, plus a recorded signature.
4.Run model = mlflow.pyfunc.load_model(model_path), then model.predict(model_input) on a sample request.
Why: This runs load_context and the full pre-process, model, post-process path the way a consumer would.
You should see: Formatted predictions come back with no import or missing-file errors.
5.Validate the logged model with mlflow.models.predict before deploying it.
Why: Databricks recommends checking that the model can be served before you deploy.
You should see: Validation completes without dependency or environment errors.
Stuck? Get a nudge
If predict fails with an import error after loading, check that the helper module was listed in code_paths and is imported inside load_context.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/machine-learning/model-serving/model-serving-custom-artifactsOfficial docs
“Real-time workloads at scale perform best when all required dependencies are statically captured at deployment time.”
↩︎ Packaging the files your chain loads“these artifacts' paths are accessible from the context object under context.artifacts”
↩︎ Packaging the files your chain loads“Model Serving requires that Unity Catalog volumes artifacts are packaged into the model artifact itself using MLflow interfaces.”
↩︎ Exam trap 2“Model Serving requires that Unity Catalog volumes artifacts are packaged into the model artifact itself using MLflow interfaces.”
↩︎ Prediction - 2.https://docs.databricks.com/aws/en/machine-learning/model-serving/deploy-custom-python-codeOfficial docs
“authors of models can log full code references that load into the path”
↩︎ Shipping shared preprocessing code“Model Serving requires mlflow (not mlflow-skinny) in conda.yaml and cannot build the container image otherwise.”
↩︎ Pinning dependencies so serving can build the image“use mlflow.models.predict to validate models before deployment”
↩︎ Signature, input example, and testing the chain before serving“you can register it to Unity Catalog or Workspace Registry and serve your model to a Model Serving endpoint”
↩︎ Signature, input example, and testing the chain before serving“Always specify mlflow==<version> in pip_requirements when you call mlflow.pyfunc.log_model() on a Databricks Runtime ML runtime”
↩︎ Exam trap 1 - 3.
“When loading the model, MLflow adds these files or directories to the Python path.”
↩︎ Shipping shared preprocessing code“MLflow infers the dependencies using mlflow.models.infer_pip_requirements, and logs them to a requirements.txt file as a model artifact.”
↩︎ Pinning dependencies so serving can build the image“you can specify additional dependencies with the extra_pip_requirements parameter in the log_model command.”
↩︎ Pinning dependencies so serving can build the image“doing so is generally discouraged because this overrides the dependencies which MLflow picks up automatically.”
↩︎ Pinning dependencies so serving can build the image“using the code_paths parameter (or code_path in MLflow 2.x)”
↩︎ Exam trap 3“using the code_paths parameter (or code_path in MLflow 2.x)”
↩︎ Checkpoint - 4.
“Adding a signature and input example to MLflow is recommended.”
↩︎ Signature, input example, and testing the chain before serving“Signatures are necessary for logging models to the Unity Catalog.”
↩︎ Checkpoint - 5.https://docs.databricks.com/aws/en/mlflow/modelsOfficial docs
“For Python MLflow models, an additional option is to use mlflow.pyfunc.load_model() to load the model as a generic Python function.”
↩︎ Signature, input example, and testing the chain before serving