CertSafari
    Databricks Certified Associate Developer for Apache Spark· Lessons

    Domain 3 · Lesson 19/32

    Python UDFs and pandas UDFs in PySpark: create, register, invoke

    Create and invoke user-defined functions with or without stateful operators, including StateStores.

    9 min read
    3.12% of exam
    3 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Create a scalar Python UDF with udf(), the @udf decorator, or spark.udf.register, and invoke it from the DataFrame API and from SQL
    • Set an explicit return type, and explain what happens when you leave it out
    • Make a UDF null-safe using the two patterns Databricks recommends
    • Write Series-to-Series and iterator pandas UDFs, and explain when an iterator UDF suits work that needs one-time setup state

    Key concept

    User-defined function (UDF) — A UDF is your own function that Spark calls inside a query, either on single values (Python UDF) or on Arrow-backed batches (pandas UDF). A UDF on its own sees only the data passed to it. Keeping state per key across streaming micro-batches needs a separate stateful operator backed by a state store.

    1.Creating and invoking a scalar Python UDF

    A scalar Python UDF starts as an ordinary Python function. Spark gives you two ways to make it callable. spark.udf.register(name, func, returnType) registers it under a name that SQL can use. pyspark.sql.functions.udf(func, returnType) wraps it as a column function for the DataFrame API. A third form is the @udf("long") decorator, which does the same job as udf() with annotation syntax. All three take the same optional return type, and that argument is where candidates slip.

    Registering a UDF for SQL with an explicit LongType return typepython
    from pyspark.sql.types import LongType
    def squared_typed(s):
      return s * s
    spark.udf.register("squaredWithPython", squared_typed, LongType())

    After registration, you call the function by name in SQL, for example select id, squaredWithPython(id) as id_squared from test. For the DataFrame API you don't need a name. You use the wrapped object directly as a column expression inside select or withColumn, as below. The return type can be given as a type object (LongType()) or as a DDL string ("long").

    Declaring the same UDF with the @udf decorator and invoking it in selectpython
    from pyspark.sql.functions import udf
    
    @udf("long")
    def squared_udf(s):
      return s * s
    df = spark.table("test")
    display(df.select("id", squared_udf("id").alias("id_squared")))

    Return types aren't limited to scalars. The Databricks examples declare returnType=StructType([...]), ArrayType(...) and MapType(StringType(), ...), and return a Python dict or list that matches the declared shape.

    Checkpoint 1 of 5· Check yourself

    A colleague wants to call a Python function from a %sql cell as myFunc(col). Which step is required?

    Sources1

    2.Null handling and evaluation order

    The Databricks UDF page warns about subexpression evaluation-order caveats in Spark SQL. In practice, a WHERE s is not null sitting next to a UDF call is not a reliable guard for that UDF. The documentation gives two safe patterns. The first is to make the UDF itself null-aware. The second is to put the null check in an IF or CASE WHEN expression and call the UDF only inside the branch where the value is known to be non-null.

    A null-safe UDF, and a conditional expression that guards the callpython
    spark.udf.register("strlen_nullsafe", lambda s: len(s) if not s is None else -1, "int")
    spark.sql("select s from test1 where s is not null and strlen_nullsafe(s) > 1") # ok
    spark.sql("select s from test1 where if(s is not null, strlen(s), null) > 1")   # ok

    Checkpoint 2 of 5· Check yourself

    Which of these is one of the two approaches Databricks recommends for UDFs that might receive nulls?

    Sources1

    3.Vectorized pandas UDFs: Series to Series

    A scalar Python UDF handles one row per call. A pandas UDF (also called a vectorized UDF) uses Apache Arrow to move data and pandas to process it. Spark splits the column into batches of rows, calls your function once per batch, and concatenates the results. You declare one with pandas_udf, either as a function or as a decorator, and the Python type hints tell Spark which kind it is. For a Series-to-Series UDF, the function must return a Series of the same length as its input. It plugs into select and withColumn just like a scalar UDF.

    Checkpoint 3 of 5· Fill the gap

    Which call turns multiply_func into a vectorized UDF?

    # Declare the function and create the UDF
    def multiply_func(a: pd.Series, b: pd.Series) -> pd.Series:
        return a * b
    
    multiply =  ? (multiply_func, returnType=LongType())

    Sources2

    4.Iterator pandas UDFs: one-time setup inside a UDF

    An iterator pandas UDF receives an iterator of batches and yields an iterator of batches. The type hint is Iterator[pd.Series] -> Iterator[pd.Series], or, for several input columns, Iterator[Tuple[pandas.Series, ...]] -> Iterator[pandas.Series]. Across all batches, the total output must still be as long as the total input. The reason to use this shape is that code placed before the loop runs once and then serves every batch, such as loading a model file. Wrap the loop in try/finally so resources are released at the end.

    An iterator UDF that uses a captured value and cleans up in finallypython
    y = 1  # value captured by the UDF closure
    
    @pandas_udf("long")
    def plus_y(batch_iter: Iterator[pd.Series]) -> Iterator[pd.Series]:
        try:
            for x in batch_iter:
                yield x + y
        finally:
            pass  # release resources here, if any
    
    df.select(plus_y(col("x"))).show()
    Choosing a UDF form by its input and output
    FormHow you declare itInput → outputTypical use
    Scalar Python UDFudf(func, LongType()) / @udf("long") / spark.udf.registerOne value → one value per rowCustom logic callable from SQL or DataFrames
    Series to Series pandas UDFpandas_udf with pd.Series hintspd.Series → pd.Series of the same lengthVectorized column arithmetic
    Iterator of Series pandas UDFpandas_udf with Iterator[pd.Series] hintsIterator of batches → iterator of batchesInitializing some state once, e.g. loading a model
    Iterator of multiple Seriespandas_udf with Iterator[Tuple[pandas.Series, ...]] hintsIterator of Series tuples → iterator of SeriesSame as above, with several input columns

    The 'state' here is setup state that lives inside one UDF execution. Stateful streaming operators work differently: they keep state information for each grouping key in a state store, and that is a separate mechanism with its own API.

    Checkpoint 4 of 5· Exam question

    What is the primary technical reason a `pandas_udf` runs faster than an equivalent standard Python `udf` when processing large DataFrames?

    Checkpoint 5 of 5· Check yourself

    You need to load a large ML model once and score every batch of a column with it. Which UDF form fits best?

    Sources23

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A Python UDF that returns numbers produces a numeric column even without a declared return type.Why is that wrong?

      Spark does not infer the type. With no return type, a Python UDF's result is StringType, so pass LongType() or "long" to get numbers back.

      Covered in Creating and invoking a scalar Python UDF

    2. 2.Because an iterator pandas UDF yields batches freely, it can return fewer rows than it received, for example to filter.Why is that wrong?

      Across the whole iterator, the output must be exactly as long as the input. Iterator UDFs are a column transformation, not a filter.

      Covered in Iterator pandas UDFs: one-time setup inside a UDF

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “You can optionally set the return type of your UDF. The default return type is StringType.”
      ↩︎ Creating and invoking a scalar Python UDF
      “Alternatively, you can declare the same UDF using annotation syntax:”
      ↩︎ Creating and invoking a scalar Python UDF
      “handle subexpression evaluation-order caveats in Spark SQL”
      ↩︎ Null handling and evaluation order
      “Make the UDF itself null-aware and do null checking inside the UDF itself”
      ↩︎ Null handling and evaluation order
      “Python scalar UDFs let you run custom Python logic inside SQL queries on Databricks.”
      ↩︎ Key concept
      “The default return type is StringType.”
      ↩︎ Exam trap 1
      “spark.udf.register("squaredWithPython", squared)”
      ↩︎ Checkpoint
      “Use IF or CASE WHEN expressions to do the null check and invoke the UDF in a conditional branch”
      ↩︎ Checkpoint
    2. 2.
      “pandas UDFs allow vectorized operations that can increase performance up to 100x compared to row-at-a-time Python UDFs.”
      ↩︎ Vectorized pandas UDFs: Series to Series
      “The Python function must accept a pandas Series as an input and return a pandas Series of the same length.”
      ↩︎ Vectorized pandas UDFs: Series to Series
      “You should specify the Python type hint as Iterator[pandas.Series] -> Iterator[pandas.Series].”
      ↩︎ Iterator pandas UDFs: one-time setup inside a UDF
      “The length of the entire output in the iterator should be the same as the length of the entire input.”
      ↩︎ Exam trap 2
      “useful when the UDF execution requires initializing some state, for example, loading a machine learning model file”
      ↩︎ Checkpoint

    Continue to page 2 of 2

    Stateful operators and StateStores with transformWithState

    Spotted a mistake, or was something unclear? Tell us.