What you will be able to do
- Write a computed mean, median, or mode into a Spark DataFrame with df.na.fill, per column and by data type
- Use scikit-learn's SimpleImputer to learn fill statistics with fit and apply them with transform
- Set a per-column imputation strategy with the AutoML imputers parameter
1.Writing the statistic back with df.na.fill
Imputing in Spark takes two steps. First you compute the statistic, for example with sf.mean, sf.median or sf.mode. Then you replace the nulls with it. DataFrame.fillna and DataFrameNaFunctions.fill, which you reach through df.na.fill, are aliases of each other, so they behave the same way. Both return a new DataFrame. Look at what value accepts: an int, float, string, bool, or a dict. It does not take a column expression. So you can't pass sf.median(...) straight into fill. You compute the median as a number first, then pass that number.
A table usually needs a different statistic for each column, such as a median for a skewed numeric column and the mode for a category. The dict form handles this in one call. You map each column name to its own replacement value, and columns that aren't in the dict keep their nulls.
df.na.fill({'age': 50, 'name': 'unknown'}).show()
# +---+------+-------+----+
# |age|height| name|bool|
# +---+------+-------+----+
# | 10| 80.5| Alice|NULL|
# | 5| NULL| Bob|NULL|
# | 50| NULL| Tom|NULL|
# | 50| NULL|unknown|true|
# +---+------+-------+----+The subset parameter limits a scalar fill to the columns you name. There are two rules to remember. First, any column in subset whose data type doesn't match the value is skipped without an error. Second, if value is a dict, subset is ignored entirely, because the dict already names the columns.
Checkpoint 1 of 6· Fill the gap
Which keyword limits this fill to the name column?
df.na.fill(value='Spark', ? ='name').show()fill(value, subset=None) takes the columns to consider in subset, given as a str, tuple, or list.
Checkpoint 2 of 6· Check yourself
You run df.na.fill({'age': 31}, subset=['age', 'height']). What happens to nulls in height?
When value is a dict, Spark uses only the dict's column-to-value mapping and ignores subset. height isn't in the dict, so its nulls stay.
“If the value is a dict, then subset is ignored and value must be a mapping from column name (string) to replacement value.”Source: docs.databricks.com
2.SimpleImputer: learn the statistic, then apply it
On pandas or NumPy data, scikit-learn's SimpleImputer does both steps in one object. The strategy argument chooses the statistic: 'mean', 'median' or 'most_frequent'. missing_values tells the imputer which placeholder means missing, such as np.nan. fit learns one statistic per column, and transform writes those learned values into whatever data you pass it.
>>> import numpy as np
>>> from sklearn.impute import SimpleImputer
>>> imp = SimpleImputer(missing_values=np.nan, strategy='mean')
>>> imp.fit([[1, 2], [np.nan, 3], [7, 6]])
SimpleImputer()
>>> X = [[np.nan, 2], [6, np.nan], [7, 6]]
>>> print(imp.transform(X))It is 4.0, the mean of 1 and 7 from the data passed to fit, not the 6.5 you would get from X. In the same way, X[1][1] becomes about 3.666, the mean of 2, 3 and 6 from the fit data. The statistic is fixed when you call fit. That is why you fit on training data and reuse the same imputer on new data.
For categorical columns you change the strategy to the most frequent value, which works on string and pandas categorical data. Because the imputer is fitted, it can sit inside a scikit-learn Pipeline with the model. One behaviour to know: by default the scikit-learn imputers drop any feature that is entirely missing, and that changes the number of columns in the output.
Checkpoint 3 of 6· Fill the gap
Which strategy makes this imputer fill each categorical column with its most common category?
>>> imp = SimpleImputer(strategy=" ? ")
>>> print(imp.fit_transform(df))scikit-learn calls the mode 'most_frequent'. 'mode' is not a valid strategy name, and 'mean' and 'median' can't handle string categories.
Source: scikit-learn.orgCheckpoint 4 of 6· Exam question
A data scientist has converted a Spark DataFrame to pandas and needs to fill missing values in a categorical column, `payment_method`, before fitting a scikit-learn model. The team wants an approach that plugs into a scikit-learn `Pipeline` and fills gaps with the most common category. Which approach correctly satisfies this requirement?
Correct answer: A — Use `SimpleImputer(strategy='most_frequent')` on the column, since it computes the mode of the observed categories and works with non-numeric string values inside a `Pipeline`.
- A. `SimpleImputer(strategy='most_frequent')` explicitly supports string and categorical dtypes and fills gaps with the mode, and it is a standard scikit-learn transformer that composes cleanly inside a `Pipeline`.
- B. `pyspark.ml.feature.Imputer` is a Spark ML transformer built for numeric input columns and Spark DataFrames; it does not accept pandas categorical columns and cannot be dropped into a scikit-learn `Pipeline`.
- C. The median strategy requires numeric input and would raise an error or produce a meaningless value on a string column, since category labels have no ordering to compute a median over.
- D. The mean strategy is only defined for numeric data; casting a categorical column to strings and averaging it does not compute a mode and would fail rather than fill the column with the most common label.
3.Setting the strategy in Databricks AutoML
Databricks AutoML lets you make the same choice without writing fill code. The imputers parameter of databricks.automl.classify and databricks.automl.regress is a dictionary. Each key is a column name, and each value is a string or dictionary describing the imputation strategy. The allowed strategy strings are the scikit-learn names mean, median and most_frequent. To fill with a known value, you pass {"strategy": "constant", "fill_value": ...}. In the UI, the same option is the Impute with column in the table schema.
If you don't specify a strategy for a column, AutoML picks a default based on the column's type and content. Setting one yourself has a side effect: AutoML does not run semantic type detection on any column that has a non-default imputation method.
Checkpoint 5 of 6· Exam question
A team needs to impute missing values in a numeric `sensor_reading` column of a multi-billion-row Spark DataFrame using the median, without collecting the full column to the driver. Which approach correctly fills the missing values at this scale, and what should the team understand about the result?
Correct answer: A — Use `pyspark.ml.feature.Imputer` with `strategy='median'`; it computes the median across executors with `approxQuantile`, so the filled value is a close approximation rather than an exact median.
- A. `pyspark.ml.feature.Imputer` scales to distributed data and, for the median strategy, internally relies on `approxQuantile` with a small relative error, so the imputed value is a close but approximate median rather than an exact one.
- B. Calling `.toPandas()` on a multi-billion-row DataFrame pulls the entire dataset into driver memory, which is exactly what the team wants to avoid at this scale, even though the resulting median would be exact.
- C. The mean strategy is also computed at this scale, but the claim that it is the only strategy guaranteeing an exact statistic is incorrect; a distributed sum-and-count mean is exact, while it is the median that is approximated, not the mean that fails.
- D. `dbutils.data.summarize()` is a diagnostic utility for displaying descriptive statistics in a notebook; it is not built to hand its output into an imputation pipeline, and manually wiring it up is not the standard distributed approach.
Checkpoint 6 of 6· Check yourself
You call AutoML with imputers={"income": "median"} and give no entry for age. How is age imputed?
imputers is set per column. Any column without an entry gets a default strategy that AutoML chooses for it.
“If no imputation strategy is provided for a column, AutoML selects a default strategy based on column type and content.”Source: docs.databricks.com
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Passing subset together with a dict value restricts the fill to the subset columns.Why is that wrong?
When value is a dict, Spark ignores subset and uses only the dict's column-to-value mapping.
Covered in Writing the statistic back with df.na.fill
2.Listing a string column in subset while filling with a numeric median raises an error, or turns the median into a string and fills it in.Why is that wrong?
Columns in subset whose data type doesn't match the value are skipped, so their nulls remain and no error is raised.
Covered in Writing the statistic back with df.na.fill
3.SimpleImputer always returns the same number of columns it was given.Why is that wrong?
By default, scikit-learn imputers drop features that are entirely missing, which changes the shape of the output.
Covered in SimpleImputer: learn the statistic, then apply it
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“DataFrame.fillna and DataFrameNaFunctions.fill are aliases of each other.”
↩︎ Writing the statistic back with df.na.fill“Fill all null values with 50 for numeric columns.”
↩︎ Prediction - 2.
“The replacement value must be an int, float, boolean, or string.”
↩︎ Writing the statistic back with df.na.fill“Columns specified in subset that do not have matching data types are ignored.”
↩︎ Writing the statistic back with df.na.fill“If the value is a dict, then subset is ignored and value must be a mapping from column name (string) to replacement value.”
↩︎ Exam trap 1“Columns specified in subset that do not have matching data types are ignored.”
↩︎ Exam trap 2“If the value is a dict, then subset is ignored and value must be a mapping from column name (string) to replacement value.”
↩︎ Checkpoint - 3.
“each value is a string or dictionary describing the imputation strategy”
↩︎ SimpleImputer: learn the statistic, then apply it“If no imputation strategy is provided for a column, AutoML selects a default strategy based on column type and content.”
↩︎ Setting the strategy in Databricks AutoML - 4.https://scikit-learn.org/stable/modules/impute.htmlSecondary source
“By default, the scikit-learn imputers will drop fully empty features, i.e. columns containing only missing values.”
↩︎ SimpleImputer: learn the statistic, then apply it“This class also allows for different missing values encodings.”
↩︎ SimpleImputer: learn the statistic, then apply it“By default, the scikit-learn imputers will drop fully empty features, i.e. columns containing only missing values.”
↩︎ Exam trap 3 - 5.
“If you specify a non-default imputation method, AutoML does not perform semantic type detection.”
↩︎ Setting the strategy in Databricks AutoML