What you will be able to do
- Create a scalar Python UDF with udf(), the @udf decorator, or spark.udf.register, and invoke it from the DataFrame API and from SQL
- Set an explicit return type, and explain what happens when you leave it out
- Make a UDF null-safe using the two patterns Databricks recommends
- Write Series-to-Series and iterator pandas UDFs, and explain when an iterator UDF suits work that needs one-time setup state
Key concept
User-defined function (UDF) — A UDF is your own function that Spark calls inside a query, either on single values (Python UDF) or on Arrow-backed batches (pandas UDF). A UDF on its own sees only the data passed to it. Keeping state per key across streaming micro-batches needs a separate stateful operator backed by a state store.
1.Creating and invoking a scalar Python UDF
A scalar Python UDF starts as an ordinary Python function. Spark gives you two ways to make it callable. spark.udf.register(name, func, returnType) registers it under a name that SQL can use. pyspark.sql.functions.udf(func, returnType) wraps it as a column function for the DataFrame API. A third form is the @udf("long") decorator, which does the same job as udf() with annotation syntax. All three take the same optional return type, and that argument is where candidates slip.
from pyspark.sql.types import LongType
def squared_typed(s):
return s * s
spark.udf.register("squaredWithPython", squared_typed, LongType())After registration, you call the function by name in SQL, for example select id, squaredWithPython(id) as id_squared from test. For the DataFrame API you don't need a name. You use the wrapped object directly as a column expression inside select or withColumn, as below. The return type can be given as a type object (LongType()) or as a DDL string ("long").
from pyspark.sql.functions import udf
@udf("long")
def squared_udf(s):
return s * s
df = spark.table("test")
display(df.select("id", squared_udf("id").alias("id_squared")))Return types aren't limited to scalars. The Databricks examples declare returnType=StructType([...]), ArrayType(...) and MapType(StringType(), ...), and return a Python dict or list that matches the declared shape.
Checkpoint 1 of 5· Check yourself
A colleague wants to call a Python function from a %sql cell as myFunc(col). Which step is required?
Only spark.udf.register gives the function a name that SQL can resolve. The udf() wrapper and the @udf decorator create a column function for the DataFrame API.
“spark.udf.register("squaredWithPython", squared)”Source: docs.databricks.com
Sources1
2.Null handling and evaluation order
The Databricks UDF page warns about subexpression evaluation-order caveats in Spark SQL. In practice, a WHERE s is not null sitting next to a UDF call is not a reliable guard for that UDF. The documentation gives two safe patterns. The first is to make the UDF itself null-aware. The second is to put the null check in an IF or CASE WHEN expression and call the UDF only inside the branch where the value is known to be non-null.
spark.udf.register("strlen_nullsafe", lambda s: len(s) if not s is None else -1, "int")
spark.sql("select s from test1 where s is not null and strlen_nullsafe(s) > 1") # ok
spark.sql("select s from test1 where if(s is not null, strlen(s), null) > 1") # okCheckpoint 2 of 5· Check yourself
Which of these is one of the two approaches Databricks recommends for UDFs that might receive nulls?
The guidance is either to make the UDF null-aware or to call it only inside a conditional branch. A plain AND predicate doesn't guarantee that the null check runs first.
“Use IF or CASE WHEN expressions to do the null check and invoke the UDF in a conditional branch”Source: docs.databricks.com
Sources1
3.Vectorized pandas UDFs: Series to Series
A scalar Python UDF handles one row per call. A pandas UDF (also called a vectorized UDF) uses Apache Arrow to move data and pandas to process it. Spark splits the column into batches of rows, calls your function once per batch, and concatenates the results. You declare one with pandas_udf, either as a function or as a decorator, and the Python type hints tell Spark which kind it is. For a Series-to-Series UDF, the function must return a Series of the same length as its input. It plugs into select and withColumn just like a scalar UDF.
Checkpoint 3 of 5· Fill the gap
Which call turns multiply_func into a vectorized UDF?
# Declare the function and create the UDF
def multiply_func(a: pd.Series, b: pd.Series) -> pd.Series:
return a * b
multiply = ? (multiply_func, returnType=LongType())pandas_udf creates the Arrow-backed vectorized UDF. Plain udf would call multiply_func one row at a time with scalar values, not Series.
Source: docs.databricks.comBecause it runs pandas operations on whole batches instead of calling Python once per row. Databricks says this can increase performance up to 100x compared to row-at-a-time Python UDFs.
Sources2
4.Iterator pandas UDFs: one-time setup inside a UDF
An iterator pandas UDF receives an iterator of batches and yields an iterator of batches. The type hint is Iterator[pd.Series] -> Iterator[pd.Series], or, for several input columns, Iterator[Tuple[pandas.Series, ...]] -> Iterator[pandas.Series]. Across all batches, the total output must still be as long as the total input. The reason to use this shape is that code placed before the loop runs once and then serves every batch, such as loading a model file. Wrap the loop in try/finally so resources are released at the end.
y = 1 # value captured by the UDF closure
@pandas_udf("long")
def plus_y(batch_iter: Iterator[pd.Series]) -> Iterator[pd.Series]:
try:
for x in batch_iter:
yield x + y
finally:
pass # release resources here, if any
df.select(plus_y(col("x"))).show()| Form | How you declare it | Input → output | Typical use |
|---|---|---|---|
| Scalar Python UDF | udf(func, LongType()) / @udf("long") / spark.udf.register | One value → one value per row | Custom logic callable from SQL or DataFrames |
| Series to Series pandas UDF | pandas_udf with pd.Series hints | pd.Series → pd.Series of the same length | Vectorized column arithmetic |
| Iterator of Series pandas UDF | pandas_udf with Iterator[pd.Series] hints | Iterator of batches → iterator of batches | Initializing some state once, e.g. loading a model |
| Iterator of multiple Series | pandas_udf with Iterator[Tuple[pandas.Series, ...]] hints | Iterator of Series tuples → iterator of Series | Same as above, with several input columns |
The 'state' here is setup state that lives inside one UDF execution. Stateful streaming operators work differently: they keep state information for each grouping key in a state store, and that is a separate mechanism with its own API.
Checkpoint 4 of 5· Exam question
What is the primary technical reason a `pandas_udf` runs faster than an equivalent standard Python `udf` when processing large DataFrames?
Correct answer: A — It transfers batches between the JVM and Python process using Arrow's columnar format, so the function runs as one vectorized pandas operation instead of many per-row calls.
- A. Arrow-based batching is the actual mechanism: rows are grouped into pandas Series or DataFrames and handed to the function as whole columns, so the function executes once per batch instead of once per row, which is where the speedup comes from.
- B. Vectorized UDFs still run as genuine Python code inside a separate Python process; nothing compiles the function body into JVM bytecode, and the Python interpreter remains fully on the execution path.
- C. There is no built-in output cache for vectorized UDFs; every partition or micro-batch still invokes the function against whatever rows it receives, with no memoization of prior results.
- D. Catalyst can push down certain native SQL functions, but an arbitrary Python function body is never translated into a Catalyst expression, so Python execution remains unavoidable for either kind of UDF.
Checkpoint 5 of 5· Check yourself
You need to load a large ML model once and score every batch of a column with it. Which UDF form fits best?
The iterator form lets you run setup once before looping over the batches. That is exactly the model-loading case the documentation describes.
“useful when the UDF execution requires initializing some state, for example, loading a machine learning model file”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A Python UDF that returns numbers produces a numeric column even without a declared return type.Why is that wrong?
Spark does not infer the type. With no return type, a Python UDF's result is StringType, so pass LongType() or "long" to get numbers back.
Covered in Creating and invoking a scalar Python UDF
2.Because an iterator pandas UDF yields batches freely, it can return fewer rows than it received, for example to filter.Why is that wrong?
Across the whole iterator, the output must be exactly as long as the input. Iterator UDFs are a column transformation, not a filter.
Covered in Iterator pandas UDFs: one-time setup inside a UDF
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/udf/pythonOfficial docs
“You can optionally set the return type of your UDF. The default return type is StringType.”
↩︎ Creating and invoking a scalar Python UDF“Alternatively, you can declare the same UDF using annotation syntax:”
↩︎ Creating and invoking a scalar Python UDF“handle subexpression evaluation-order caveats in Spark SQL”
↩︎ Null handling and evaluation order“Make the UDF itself null-aware and do null checking inside the UDF itself”
↩︎ Null handling and evaluation order“Python scalar UDFs let you run custom Python logic inside SQL queries on Databricks.”
↩︎ Key concept“The default return type is StringType.”
↩︎ Exam trap 1“spark.udf.register("squaredWithPython", squared)”
↩︎ Checkpoint“Use IF or CASE WHEN expressions to do the null check and invoke the UDF in a conditional branch”
↩︎ Checkpoint - 2.https://docs.databricks.com/aws/en/udf/pandasOfficial docs
“pandas UDFs allow vectorized operations that can increase performance up to 100x compared to row-at-a-time Python UDFs.”
↩︎ Vectorized pandas UDFs: Series to Series“The Python function must accept a pandas Series as an input and return a pandas Series of the same length.”
↩︎ Vectorized pandas UDFs: Series to Series“You should specify the Python type hint as Iterator[pandas.Series] -> Iterator[pandas.Series].”
↩︎ Iterator pandas UDFs: one-time setup inside a UDF“The length of the entire output in the iterator should be the same as the length of the entire input.”
↩︎ Exam trap 2“useful when the UDF execution requires initializing some state, for example, loading a machine learning model file”
↩︎ Checkpoint - 3.
“State information persists for each grouping key.”
↩︎ Iterator pandas UDFs: one-time setup inside a UDF