CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 4 · Lesson 45/48

    Batch Inference with MLflow pyfunc Models and Spark

    Use pandas to perform batch inference

    12 min read
    2.08% of exam
    7 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Load a logged or registered MLflow model with mlflow.pyfunc.load_model and score data with predict
    • Turn an MLflow model into a Spark UDF with mlflow.pyfunc.spark_udf and add a prediction column with withColumn
    • Restore a model's dependencies for batch scoring with get_model_dependencies or env_manager="virtualenv"
    • Set up a batch inference pipeline that keeps preprocessing separate from scoring

    Key concept

    pandas UDF batch scoring — Batch inference applies a model to a whole table rather than one request at a time. To score a Spark table across a cluster, Databricks wraps the model as a pandas UDF: Spark cuts the input table into batches of rows, hands each batch to the model as pandas data, and joins the predictions back into one Spark DataFrame. When the data is already a pandas DataFrame, mlflow.pyfunc.load_model(...).predict scores it directly in the process that loaded the model.

    1.Why batch inference on Databricks runs through pandas

    Batch inference scores a whole table at once, usually on a schedule. It does not answer one request at a time. Databricks describes batch and streaming scoring as high-throughput, low-cost scoring at latencies as low as minutes. A typical example is scoring every customer overnight. Shaving milliseconds off a single prediction doesn't matter for that kind of job.

    Databricks' guidance is short: "Use Spark Pandas UDFs to scale batch and streaming inference across a cluster." A pandas UDF, also called a vectorized UDF, uses Apache Arrow to move data between Spark and Python and uses pandas to work on that data. It processes many rows per call instead of one, which Databricks says can make it up to 100x faster than a row-at-a-time Python UDF. This matters for inference because most model libraries are built to predict on many rows at once.

    pandas is already available. Databricks Runtime ships it as a standard Python package, so the same pandas code runs in notebooks and in jobs. In practice, you write and test scoring logic against local pandas data, then let Spark run it on every batch of the full table.

    Checkpoint 1 of 7· Check yourself

    Why do pandas UDFs beat ordinary Python UDFs for batch scoring?

    Sources123

    2.Loading the model: pyfunc predict versus spark_udf

    Every flavor of MLflow model has its own loader, mlflow.<model-type>.load_model(modelpath). For Python models there is also a generic option: mlflow.pyfunc.load_model() loads any model "as a generic Python function". Your scoring code then calls predict the same way whether the model came from scikit-learn, XGBoost or something else. The modelpath argument accepts several forms:

    - a run-relative path, such as runs:/{run_id}/{model-path} - a registered model path, such as models:/{model_name}/{model_stage} - a model path such as models:/{model_id} (MLflow 3 only) - a Unity Catalog volumes path - an MLflow-managed artifact path beginning with dbfs:/databricks/mlflow-tracking/

    Load an MLflow model as a generic Python function and score an input in a single processpython
    model = mlflow.pyfunc.load_model(model_path)
    model.predict(model_input)

    That call scores data in whatever process loaded the model. The input can be a pandas DataFrame. Databricks' AutoML examples show both routes for pandas data and Spark data: "make predictions on data in pandas DataFrames, or register the model as a Spark UDF for prediction on Spark DataFrames". If your data is a Spark DataFrame, toPandas() turns it into a pandas DataFrame, and you pass the feature columns to predict. The result can be assigned back as a new column.

    Convert a Spark DataFrame to pandas, score it in-process with load_model, and store the predictionspython
    # Prepare test dataset
    test_pdf = test_df.toPandas()
    y_test = test_pdf["income"]
    X_test = test_pdf.drop("income", axis=1)
    
    # Run inference using the best model
    model = mlflow.pyfunc.load_model(model_uri)
    predictions = model.predict(X_test)
    test_pdf["income_predicted"] = predictions
    display(test_pdf)

    To score a full table on a cluster, export the same model as a Spark UDF with mlflow.pyfunc.spark_udf(spark, model_path) and add its output as a new column. Databricks notes that the exported UDF can run "either as a batch job or as a real-time Spark Streaming job". This lesson covers the batch case. Going the other way, spark.createDataFrame(pandas_df) turns a pandas DataFrame into a Spark DataFrame so the UDF can score it.

    Batch-score a table by exporting the model as a Spark UDF and adding a prediction columnpython
    # load input data table as a Spark DataFrame
    input_data = spark.table(input_table_name)
    model_udf = mlflow.pyfunc.spark_udf(spark, model_path)
    df = input_data.withColumn("prediction", model_udf())
    Turn a pandas DataFrame into a Spark DataFrame and score it with the Spark UDFpython
    test_df = spark.createDataFrame(test_pdf)
    predict_udf = mlflow.pyfunc.spark_udf(spark, model_uri=model_uri)
    display(test_df.withColumn("MedHouseVal_predicted", predict_udf()))

    spark_udf also takes extra arguments. result_type declares the type of the prediction column, for example result_type="string" for a classifier that returns labels. env_manager controls dependency restoration, covered in the next section. When the model needs several input columns, a pipeline example from Databricks passes the feature columns to the UDF packed as a struct.

    Declare the prediction column type with result_typepython
    predict_udf = mlflow.pyfunc.spark_udf(spark, model_uri=model_uri, result_type="string")
    display(test_df.withColumn("income_predicted", predict_udf()))
    Pass the feature columns to the model UDF as a structpython
    .withColumn('predictions', loaded_model_udf(struct(features)))
    Two ways to apply the same MLflow model
    CallWhat you getWhere scoring runs
    mlflow.pyfunc.load_model(model_path)A generic Python function model with a predict methodIn the process that loaded it, via model.predict(model_input)
    mlflow.pyfunc.spark_udf(spark, model_path)A Spark UDF you apply as a column expressionAcross the Spark cluster, via withColumn("prediction", model_udf())

    Checkpoint 2 of 7· Fill the gap

    Which MLflow function completes this cluster-wide batch scoring sample?

    # load input data table as a Spark DataFrame
    input_data = spark.table(input_table_name)
    model_udf = mlflow.pyfunc. ? (spark, model_path)
    df = input_data.withColumn("prediction", model_udf())

    Checkpoint 3 of 7· Check yourself

    You already hold a pandas DataFrame X_test of feature columns and a model loaded with mlflow.pyfunc.load_model. How do the AutoML examples score it?

    Checkpoint 4 of 7· Exam question

    A data engineer has a Delta table with 50 million transaction rows and a model registered at `models:/fraud_detector/Production`. They need to score every row and write predictions to a new Delta table while distributing the scoring work across the cluster. Which approach should they use?

    Checkpoint 5 of 7· Exam question

    A data scientist has a pandas DataFrame with 10,000 rows already loaded in the driver's memory and a model registered at `models:/churn_model/Production`. They want to generate predictions for this one-off analysis with the least setup overhead. Which approach best fits this situation?

    Sources456

    3.Making the scoring environment match the model

    A model loaded with the wrong library versions can fail, or it can quietly give different predictions. Databricks says that to load a model accurately, you should make sure its dependencies are installed at the correct versions in the notebook environment. Databricks Runtime gives you three aids, depending on version:

    - Databricks Runtime 10.5 ML and above: MLflow warns you when the current environment doesn't match the model's dependencies. - Databricks Runtime 11.0 ML and above, notebook scoring: for pyfunc-flavor models, mlflow.pyfunc.get_model_dependencies downloads the model's dependency file and returns its path. You then install it with %pip install <file-path>. - Databricks Runtime 11.0 ML and above, Spark UDF scoring: pass env_manager="virtualenv" to mlflow.pyfunc.spark_udf. The dependencies are restored inside the UDF, and the notebook's own environment is left alone.

    Checkpoint 6 of 7· Check yourself

    What happens to the notebook's own Python environment when you pass env_manager="virtualenv" to mlflow.pyfunc.spark_udf?

    Sources4

    4.Designing the batch pipeline around the model

    The scoring call is only one step of the job. If the same data will be scored more than once, Databricks recommends a separate preprocessing job that ETLs it into a Delta Lake table first. The cost of ingesting and preparing the data is then paid once and shared across every inference run. The split also lets you choose hardware for each job, for example CPUs for ETL and GPUs for inference.

    The Databricks reference solution for image models uses the same two stages. First, Auto Loader ETLs images into a Delta table and picks up new images as they arrive. Then a pandas UDF runs distributed inference over that table. The guide calls the inference "embarrassingly parallel": each batch can be scored independently, which is exactly the case pandas UDFs are built for. For very large images (over 100 MB on average), the Delta table holds only the file names, and images are loaded from object storage by path when needed.

    Checkpoint 7 of 7· Check yourself

    A team rescores the same raw clickstream data every week with a new model version. Each run spends most of its time parsing and cleaning raw files before scoring. What does Databricks recommend?

    Sources17

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Restoring a model's dependencies for spark_udf scoring overwrites the notebook's installed library versions.Why is that wrong?

      With env_manager="virtualenv", the dependencies are restored only inside the PySpark UDF. The rest of the notebook keeps its own library versions.

      Covered in Making the scoring environment match the model

    2. 2.Scaling batch scoring across a cluster requires deploying the model to an online serving endpoint.Why is that wrong?

      Databricks scales batch scoring with Spark pandas UDFs applied to the data. Online serving is a separate deployment mode for low-latency requests.

      Covered in Why batch inference on Databricks runs through pandas

    3. 3.A model loaded with mlflow.pyfunc.load_model can only be used through a Spark UDF, never on a pandas DataFrame.Why is that wrong?

      The AutoML examples call model.predict on a pandas DataFrame directly. spark_udf is the route for scoring Spark DataFrames on the cluster.

      Covered in Loading the model: pyfunc predict versus spark_udf

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Batch and streaming scoring supports high-throughput, low-cost scoring at latencies as low as minutes.”
      ↩︎ Why batch inference on Databricks runs through pandas
      “When you log a model from Databricks, MLflow automatically provides inference code to apply the model as a pandas UDF.”
      ↩︎ Why batch inference on Databricks runs through pandas
      “you might use CPUs for ETL and GPUs for inference.”
      ↩︎ Designing the batch pipeline around the model
      “Use Spark Pandas UDFs to scale batch and streaming inference across a cluster.”
      ↩︎ Exam trap 2
      “consider creating a preprocessing job to ETL the data into a Delta Lake table before running the inference job.”
      ↩︎ Checkpoint
    2. 2.
      “uses Apache Arrow to transfer data and pandas to work with the data”
      ↩︎ Why batch inference on Databricks runs through pandas
      “Spark runs a pandas UDF by splitting the data into batches of rows, calling the function for each batch, and then concatenating the results.”
      ↩︎ Key concept
      “pandas UDFs allow vectorized operations that can increase performance up to 100x compared to row-at-a-time Python UDFs.”
      ↩︎ Checkpoint
    3. 3.
      “Databricks Runtime includes pandas as one of the standard Python packages”
      ↩︎ Why batch inference on Databricks runs through pandas
    4. 4.
      “For Python MLflow models, an additional option is to use mlflow.pyfunc.load_model() to load the model as a generic Python function.”
      ↩︎ Loading the model: pyfunc predict versus spark_udf
      “export the model as an Apache Spark UDF to use for scoring on a Spark cluster”
      ↩︎ Loading the model: pyfunc predict versus spark_udf
      “a registered model path (such as models:/{model_name}/{model_stage}).”
      ↩︎ Loading the model: pyfunc predict versus spark_udf
      “In Databricks Runtime 10.5 ML and above, MLflow warns you if a mismatch is detected between the current environment and the model's dependencies.”
      ↩︎ Making the scoring environment match the model
      “you can call mlflow.pyfunc.get_model_dependencies to retrieve and download the model dependencies.”
      ↩︎ Making the scoring environment match the model
      “This restores model dependencies in the context of the PySpark UDF and does not affect the outside environment.”
      ↩︎ Exam trap 1
      “This restores model dependencies in the context of the PySpark UDF and does not affect the outside environment.”
      ↩︎ Checkpoint
    5. 5.
      “make predictions on data in pandas DataFrames, or register the model as a Spark UDF for prediction on Spark DataFrames”
      ↩︎ Loading the model: pyfunc predict versus spark_udf
      “make predictions on data in pandas DataFrames, or register the model as a Spark UDF for prediction on Spark DataFrames”
      ↩︎ Exam trap 3
    6. 6.
      “The data columns used to make the prediction are passed as an argument to the UDF.”
      ↩︎ Loading the model: pyfunc predict versus spark_udf
    7. 7.
      “the inference workload is embarrassingly parallel and in theory can be distributed easily.”
      ↩︎ Designing the batch pipeline around the model

    Continue to page 2 of 2

    pandas UDF Types and Function APIs for Model Scoring

    Spotted a mistake, or was something unclear? Tell us.