What you will be able to do
- Load a logged or registered MLflow model with mlflow.pyfunc.load_model and score data with predict
- Turn an MLflow model into a Spark UDF with mlflow.pyfunc.spark_udf and add a prediction column with withColumn
- Restore a model's dependencies for batch scoring with get_model_dependencies or env_manager="virtualenv"
- Set up a batch inference pipeline that keeps preprocessing separate from scoring
Key concept
pandas UDF batch scoring — Batch inference applies a model to a whole table rather than one request at a time. To score a Spark table across a cluster, Databricks wraps the model as a pandas UDF: Spark cuts the input table into batches of rows, hands each batch to the model as pandas data, and joins the predictions back into one Spark DataFrame. When the data is already a pandas DataFrame, mlflow.pyfunc.load_model(...).predict scores it directly in the process that loaded the model.
1.Why batch inference on Databricks runs through pandas
Batch inference scores a whole table at once, usually on a schedule. It does not answer one request at a time. Databricks describes batch and streaming scoring as high-throughput, low-cost scoring at latencies as low as minutes. A typical example is scoring every customer overnight. Shaving milliseconds off a single prediction doesn't matter for that kind of job.
Databricks' guidance is short: "Use Spark Pandas UDFs to scale batch and streaming inference across a cluster." A pandas UDF, also called a vectorized UDF, uses Apache Arrow to move data between Spark and Python and uses pandas to work on that data. It processes many rows per call instead of one, which Databricks says can make it up to 100x faster than a row-at-a-time Python UDF. This matters for inference because most model libraries are built to predict on many rows at once.
pandas is already available. Databricks Runtime ships it as a standard Python package, so the same pandas code runs in notebooks and in jobs. In practice, you write and test scoring logic against local pandas data, then let Spark run it on every batch of the full table.
Checkpoint 1 of 7· Check yourself
Why do pandas UDFs beat ordinary Python UDFs for batch scoring?
The speedup comes from Arrow-based transfer plus vectorized pandas operations over whole batches, compared with calling Python once per row.
“pandas UDFs allow vectorized operations that can increase performance up to 100x compared to row-at-a-time Python UDFs.”Source: docs.databricks.com
2.Loading the model: pyfunc predict versus spark_udf
Every flavor of MLflow model has its own loader, mlflow.<model-type>.load_model(modelpath). For Python models there is also a generic option: mlflow.pyfunc.load_model() loads any model "as a generic Python function". Your scoring code then calls predict the same way whether the model came from scikit-learn, XGBoost or something else. The modelpath argument accepts several forms:
- a run-relative path, such as runs:/{run_id}/{model-path}
- a registered model path, such as models:/{model_name}/{model_stage}
- a model path such as models:/{model_id} (MLflow 3 only)
- a Unity Catalog volumes path
- an MLflow-managed artifact path beginning with dbfs:/databricks/mlflow-tracking/
model = mlflow.pyfunc.load_model(model_path)
model.predict(model_input)That call scores data in whatever process loaded the model. The input can be a pandas DataFrame. Databricks' AutoML examples show both routes for pandas data and Spark data: "make predictions on data in pandas DataFrames, or register the model as a Spark UDF for prediction on Spark DataFrames". If your data is a Spark DataFrame, toPandas() turns it into a pandas DataFrame, and you pass the feature columns to predict. The result can be assigned back as a new column.
# Prepare test dataset
test_pdf = test_df.toPandas()
y_test = test_pdf["income"]
X_test = test_pdf.drop("income", axis=1)
# Run inference using the best model
model = mlflow.pyfunc.load_model(model_uri)
predictions = model.predict(X_test)
test_pdf["income_predicted"] = predictions
display(test_pdf)To score a full table on a cluster, export the same model as a Spark UDF with mlflow.pyfunc.spark_udf(spark, model_path) and add its output as a new column. Databricks notes that the exported UDF can run "either as a batch job or as a real-time Spark Streaming job". This lesson covers the batch case. Going the other way, spark.createDataFrame(pandas_df) turns a pandas DataFrame into a Spark DataFrame so the UDF can score it.
# load input data table as a Spark DataFrame
input_data = spark.table(input_table_name)
model_udf = mlflow.pyfunc.spark_udf(spark, model_path)
df = input_data.withColumn("prediction", model_udf())test_df = spark.createDataFrame(test_pdf)
predict_udf = mlflow.pyfunc.spark_udf(spark, model_uri=model_uri)
display(test_df.withColumn("MedHouseVal_predicted", predict_udf()))spark_udf also takes extra arguments. result_type declares the type of the prediction column, for example result_type="string" for a classifier that returns labels. env_manager controls dependency restoration, covered in the next section. When the model needs several input columns, a pipeline example from Databricks passes the feature columns to the UDF packed as a struct.
predict_udf = mlflow.pyfunc.spark_udf(spark, model_uri=model_uri, result_type="string")
display(test_df.withColumn("income_predicted", predict_udf())).withColumn('predictions', loaded_model_udf(struct(features)))| Call | What you get | Where scoring runs |
|---|---|---|
| mlflow.pyfunc.load_model(model_path) | A generic Python function model with a predict method | In the process that loaded it, via model.predict(model_input) |
| mlflow.pyfunc.spark_udf(spark, model_path) | A Spark UDF you apply as a column expression | Across the Spark cluster, via withColumn("prediction", model_udf()) |
Checkpoint 2 of 7· Fill the gap
Which MLflow function completes this cluster-wide batch scoring sample?
# load input data table as a Spark DataFrame
input_data = spark.table(input_table_name)
model_udf = mlflow.pyfunc. ? (spark, model_path)
df = input_data.withColumn("prediction", model_udf())mlflow.pyfunc.spark_udf takes the SparkSession and the model path and returns a UDF that withColumn can apply. load_model returns a model object for in-process predict, not a column expression.
Checkpoint 3 of 7· Check yourself
You already hold a pandas DataFrame X_test of feature columns and a model loaded with mlflow.pyfunc.load_model. How do the AutoML examples score it?
The AutoML examples pass pandas data straight to model.predict, and use spark_udf when the data is a Spark DataFrame.
“make predictions on data in pandas DataFrames, or register the model as a Spark UDF for prediction on Spark DataFrames”Source: docs.databricks.com
Checkpoint 4 of 7· Exam question
A data engineer has a Delta table with 50 million transaction rows and a model registered at `models:/fraud_detector/Production`. They need to score every row and write predictions to a new Delta table while distributing the scoring work across the cluster. Which approach should they use?
Correct answer: B — Load the model with `mlflow.pyfunc.spark_udf(spark, "models:/fraud_detector/Production")` and add predictions to the Spark DataFrame via `withColumn`, letting executors score partitions in parallel.
- A. Collecting 50 million rows into a single pandas DataFrame with `.toPandas()` pulls the entire dataset into driver memory, which risks an out-of-memory failure and abandons the cluster's distributed executors entirely. It also removes any parallelism, since one `.predict()` call runs sequentially on the driver.
- B. Wrapping the registered model as a pandas UDF with `mlflow.pyfunc.spark_udf` and applying it through `withColumn` lets each executor load the model once and score its own partition in parallel, which is the standard pattern for distributed batch inference over a large Delta table.
- C. Issuing 50 million individual REST calls from a single driver notebook is extremely slow, adds network latency per row, and does not use the cluster's compute for scoring at all, making it impractical at this scale.
- D. This SQL function can call a served model per row, but it requires a live serving endpoint to already be deployed and bypasses pandas or Spark DataFrame scoring altogether, so it does not match a pandas-based batch scoring workflow over an existing Delta table.
Checkpoint 5 of 7· Exam question
A data scientist has a pandas DataFrame with 10,000 rows already loaded in the driver's memory and a model registered at `models:/churn_model/Production`. They want to generate predictions for this one-off analysis with the least setup overhead. Which approach best fits this situation?
Correct answer: D — Call `mlflow.pyfunc.load_model("models:/churn_model/Production")` and pass the pandas DataFrame directly to `.predict()`, since the whole dataset already fits comfortably in driver memory.
- A. Converting a DataFrame that already fits in driver memory into a Spark DataFrame just to build a distributed pandas UDF adds Spark job scheduling and Arrow serialization overhead for no real benefit at this size, making it unnecessarily heavy for a 10,000-row one-off job.
- B. Standing up a real-time serving endpoint and calling it row by row from a loop introduces network round trips and deployment overhead for what is a small, local, one-time scoring task, and it is slower than scoring the DataFrame directly in process.
- C. `mlflow.pyfunc.load_model()` does not take a `stage` argument to select an alias; model stage or alias selection belongs in the model URI itself, so this call does not do what the scenario needs and adds an unnecessary new model version.
- D. Loading the model as a `PyFuncModel` and calling `.predict()` directly on the in-memory pandas DataFrame is the simplest path when the data already fits comfortably without Spark, avoiding any cluster job overhead for a small one-off scoring task.
3.Making the scoring environment match the model
A model loaded with the wrong library versions can fail, or it can quietly give different predictions. Databricks says that to load a model accurately, you should make sure its dependencies are installed at the correct versions in the notebook environment. Databricks Runtime gives you three aids, depending on version:
- Databricks Runtime 10.5 ML and above: MLflow warns you when the current environment doesn't match the model's dependencies.
- Databricks Runtime 11.0 ML and above, notebook scoring: for pyfunc-flavor models, mlflow.pyfunc.get_model_dependencies downloads the model's dependency file and returns its path. You then install it with %pip install <file-path>.
- Databricks Runtime 11.0 ML and above, Spark UDF scoring: pass env_manager="virtualenv" to mlflow.pyfunc.spark_udf. The dependencies are restored inside the UDF, and the notebook's own environment is left alone.
Pass env_manager="virtualenv" to mlflow.pyfunc.spark_udf. The model's dependencies are restored only inside the PySpark UDF, so the notebook keeps its newer scikit-learn.
Checkpoint 6 of 7· Check yourself
What happens to the notebook's own Python environment when you pass env_manager="virtualenv" to mlflow.pyfunc.spark_udf?
The virtualenv option isolates the model's dependencies inside the UDF's execution context, so code outside the UDF keeps its own library versions.
“This restores model dependencies in the context of the PySpark UDF and does not affect the outside environment.”Source: docs.databricks.com
Sources4
4.Designing the batch pipeline around the model
The scoring call is only one step of the job. If the same data will be scored more than once, Databricks recommends a separate preprocessing job that ETLs it into a Delta Lake table first. The cost of ingesting and preparing the data is then paid once and shared across every inference run. The split also lets you choose hardware for each job, for example CPUs for ETL and GPUs for inference.
The Databricks reference solution for image models uses the same two stages. First, Auto Loader ETLs images into a Delta table and picks up new images as they arrive. Then a pandas UDF runs distributed inference over that table. The guide calls the inference "embarrassingly parallel": each batch can be scored independently, which is exactly the case pandas UDFs are built for. For very large images (over 100 MB on average), the Delta table holds only the file names, and images are loaded from object storage by path when needed.
Checkpoint 7 of 7· Check yourself
A team rescores the same raw clickstream data every week with a new model version. Each run spends most of its time parsing and cleaning raw files before scoring. What does Databricks recommend?
When the data is read for inference more than once, preprocessing it into Delta spreads the cost of ingestion across all later reads. It also lets ETL and inference run on different hardware.
“consider creating a preprocessing job to ETL the data into a Delta Lake table before running the inference job.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Restoring a model's dependencies for spark_udf scoring overwrites the notebook's installed library versions.Why is that wrong?
With env_manager="virtualenv", the dependencies are restored only inside the PySpark UDF. The rest of the notebook keeps its own library versions.
2.Scaling batch scoring across a cluster requires deploying the model to an online serving endpoint.Why is that wrong?
Databricks scales batch scoring with Spark pandas UDFs applied to the data. Online serving is a separate deployment mode for low-latency requests.
Covered in Why batch inference on Databricks runs through pandas
3.A model loaded with mlflow.pyfunc.load_model can only be used through a Spark UDF, never on a pandas DataFrame.Why is that wrong?
The AutoML examples call model.predict on a pandas DataFrame directly. spark_udf is the route for scoring Spark DataFrames on the cluster.
Covered in Loading the model: pyfunc predict versus spark_udf
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Batch and streaming scoring supports high-throughput, low-cost scoring at latencies as low as minutes.”
↩︎ Why batch inference on Databricks runs through pandas“When you log a model from Databricks, MLflow automatically provides inference code to apply the model as a pandas UDF.”
↩︎ Why batch inference on Databricks runs through pandas“you might use CPUs for ETL and GPUs for inference.”
↩︎ Designing the batch pipeline around the model“Use Spark Pandas UDFs to scale batch and streaming inference across a cluster.”
↩︎ Exam trap 2“consider creating a preprocessing job to ETL the data into a Delta Lake table before running the inference job.”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/udf/pandasOfficial docs
“uses Apache Arrow to transfer data and pandas to work with the data”
↩︎ Why batch inference on Databricks runs through pandas“Spark runs a pandas UDF by splitting the data into batches of rows, calling the function for each batch, and then concatenating the results.”
↩︎ Key concept“pandas UDFs allow vectorized operations that can increase performance up to 100x compared to row-at-a-time Python UDFs.”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/pandasOfficial docs
“Databricks Runtime includes pandas as one of the standard Python packages”
↩︎ Why batch inference on Databricks runs through pandas - 4.https://docs.databricks.com/aws/en/mlflow/modelsOfficial docs
“For Python MLflow models, an additional option is to use mlflow.pyfunc.load_model() to load the model as a generic Python function.”
↩︎ Loading the model: pyfunc predict versus spark_udf“export the model as an Apache Spark UDF to use for scoring on a Spark cluster”
↩︎ Loading the model: pyfunc predict versus spark_udf“a registered model path (such as models:/{model_name}/{model_stage}).”
↩︎ Loading the model: pyfunc predict versus spark_udf“In Databricks Runtime 10.5 ML and above, MLflow warns you if a mismatch is detected between the current environment and the model's dependencies.”
↩︎ Making the scoring environment match the model“you can call mlflow.pyfunc.get_model_dependencies to retrieve and download the model dependencies.”
↩︎ Making the scoring environment match the model“This restores model dependencies in the context of the PySpark UDF and does not affect the outside environment.”
↩︎ Exam trap 1“This restores model dependencies in the context of the PySpark UDF and does not affect the outside environment.”
↩︎ Checkpoint - 5.
“make predictions on data in pandas DataFrames, or register the model as a Spark UDF for prediction on Spark DataFrames”
↩︎ Loading the model: pyfunc predict versus spark_udf“make predictions on data in pandas DataFrames, or register the model as a Spark UDF for prediction on Spark DataFrames”
↩︎ Exam trap 3 - 6.https://docs.databricks.com/aws/en/ldp/transformOfficial docs
“The data columns used to make the prediction are passed as an argument to the UDF.”
↩︎ Loading the model: pyfunc predict versus spark_udf - 7.https://docs.databricks.com/aws/en/machine-learning/reference-solutions/images-etl-inferenceOfficial docs
“the inference workload is embarrassingly parallel and in theory can be distributed easily.”
↩︎ Designing the batch pipeline around the model