CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 2 · Lesson 24/48

    Filling Missing Values with fillna, SimpleImputer, and AutoML imputers

    Impute missing values with the mode, mean, or median value

    9 min read
    2.08% of exam
    5 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Write a computed mean, median, or mode into a Spark DataFrame with df.na.fill, per column and by data type
    • Use scikit-learn's SimpleImputer to learn fill statistics with fit and apply them with transform
    • Set a per-column imputation strategy with the AutoML imputers parameter

    1.Writing the statistic back with df.na.fill

    Imputing in Spark takes two steps. First you compute the statistic, for example with sf.mean, sf.median or sf.mode. Then you replace the nulls with it. DataFrame.fillna and DataFrameNaFunctions.fill, which you reach through df.na.fill, are aliases of each other, so they behave the same way. Both return a new DataFrame. Look at what value accepts: an int, float, string, bool, or a dict. It does not take a column expression. So you can't pass sf.median(...) straight into fill. You compute the median as a number first, then pass that number.

    A table usually needs a different statistic for each column, such as a median for a skewed numeric column and the mode for a category. The dict form handles this in one call. You map each column name to its own replacement value, and columns that aren't in the dict keep their nulls.

    Per-column fill values from a dict. height is not in the dict, so it keeps its NULLspython
    df.na.fill({'age': 50, 'name': 'unknown'}).show()
    # +---+------+-------+----+
    # |age|height|   name|bool|
    # +---+------+-------+----+
    # | 10|  80.5|  Alice|NULL|
    # |  5|  NULL|    Bob|NULL|
    # | 50|  NULL|    Tom|NULL|
    # | 50|  NULL|unknown|true|
    # +---+------+-------+----+

    The subset parameter limits a scalar fill to the columns you name. There are two rules to remember. First, any column in subset whose data type doesn't match the value is skipped without an error. Second, if value is a dict, subset is ignored entirely, because the dict already names the columns.

    Checkpoint 1 of 6· Fill the gap

    Which keyword limits this fill to the name column?

    df.na.fill(value='Spark',  ? ='name').show()

    Checkpoint 2 of 6· Check yourself

    You run df.na.fill({'age': 31}, subset=['age', 'height']). What happens to nulls in height?

    Sources12

    2.SimpleImputer: learn the statistic, then apply it

    On pandas or NumPy data, scikit-learn's SimpleImputer does both steps in one object. The strategy argument chooses the statistic: 'mean', 'median' or 'most_frequent'. missing_values tells the imputer which placeholder means missing, such as np.nan. fit learns one statistic per column, and transform writes those learned values into whatever data you pass it.

    Mean imputation. fit learns column means from one dataset, and transform fills a different onepython
    >>> import numpy as np
    >>> from sklearn.impute import SimpleImputer
    >>> imp = SimpleImputer(missing_values=np.nan, strategy='mean')
    >>> imp.fit([[1, 2], [np.nan, 3], [7, 6]])
    SimpleImputer()
    >>> X = [[np.nan, 2], [6, np.nan], [7, 6]]
    >>> print(imp.transform(X))

    For categorical columns you change the strategy to the most frequent value, which works on string and pandas categorical data. Because the imputer is fitted, it can sit inside a scikit-learn Pipeline with the model. One behaviour to know: by default the scikit-learn imputers drop any feature that is entirely missing, and that changes the number of columns in the output.

    Checkpoint 3 of 6· Fill the gap

    Which strategy makes this imputer fill each categorical column with its most common category?

    >>> imp = SimpleImputer(strategy=" ? ")
    >>> print(imp.fit_transform(df))

    Checkpoint 4 of 6· Exam question

    A data scientist has converted a Spark DataFrame to pandas and needs to fill missing values in a categorical column, `payment_method`, before fitting a scikit-learn model. The team wants an approach that plugs into a scikit-learn `Pipeline` and fills gaps with the most common category. Which approach correctly satisfies this requirement?

    Sources34

    3.Setting the strategy in Databricks AutoML

    Databricks AutoML lets you make the same choice without writing fill code. The imputers parameter of databricks.automl.classify and databricks.automl.regress is a dictionary. Each key is a column name, and each value is a string or dictionary describing the imputation strategy. The allowed strategy strings are the scikit-learn names mean, median and most_frequent. To fill with a known value, you pass {"strategy": "constant", "fill_value": ...}. In the UI, the same option is the Impute with column in the table schema.

    If you don't specify a strategy for a column, AutoML picks a default based on the column's type and content. Setting one yourself has a side effect: AutoML does not run semantic type detection on any column that has a non-default imputation method.

    Checkpoint 5 of 6· Exam question

    A team needs to impute missing values in a numeric `sensor_reading` column of a multi-billion-row Spark DataFrame using the median, without collecting the full column to the driver. Which approach correctly fills the missing values at this scale, and what should the team understand about the result?

    Checkpoint 6 of 6· Check yourself

    You call AutoML with imputers={"income": "median"} and give no entry for age. How is age imputed?

    Sources35

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Passing subset together with a dict value restricts the fill to the subset columns.Why is that wrong?

      When value is a dict, Spark ignores subset and uses only the dict's column-to-value mapping.

      Covered in Writing the statistic back with df.na.fill

    2. 2.Listing a string column in subset while filling with a numeric median raises an error, or turns the median into a string and fills it in.Why is that wrong?

      Columns in subset whose data type doesn't match the value are skipped, so their nulls remain and no error is raised.

      Covered in Writing the statistic back with df.na.fill

    3. 3.SimpleImputer always returns the same number of columns it was given.Why is that wrong?

      By default, scikit-learn imputers drop features that are entirely missing, which changes the shape of the output.

      Covered in SimpleImputer: learn the statistic, then apply it

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “DataFrame.fillna and DataFrameNaFunctions.fill are aliases of each other.”
      ↩︎ Writing the statistic back with df.na.fill
      “Fill all null values with 50 for numeric columns.”
      ↩︎ Prediction
    2. 2.
      “The replacement value must be an int, float, boolean, or string.”
      ↩︎ Writing the statistic back with df.na.fill
      “Columns specified in subset that do not have matching data types are ignored.”
      ↩︎ Writing the statistic back with df.na.fill
      “If the value is a dict, then subset is ignored and value must be a mapping from column name (string) to replacement value.”
      ↩︎ Exam trap 1
      “Columns specified in subset that do not have matching data types are ignored.”
      ↩︎ Exam trap 2
      “If the value is a dict, then subset is ignored and value must be a mapping from column name (string) to replacement value.”
      ↩︎ Checkpoint
    3. 3.
      “each value is a string or dictionary describing the imputation strategy”
      ↩︎ SimpleImputer: learn the statistic, then apply it
      “If no imputation strategy is provided for a column, AutoML selects a default strategy based on column type and content.”
      ↩︎ Setting the strategy in Databricks AutoML
    4. 4.
      “By default, the scikit-learn imputers will drop fully empty features, i.e. columns containing only missing values.”
      ↩︎ SimpleImputer: learn the statistic, then apply it
      “This class also allows for different missing values encodings.”
      ↩︎ SimpleImputer: learn the statistic, then apply it
      “By default, the scikit-learn imputers will drop fully empty features, i.e. columns containing only missing values.”
      ↩︎ Exam trap 3
    5. 5.
      “If you specify a non-default imputation method, AutoML does not perform semantic type detection.”
      ↩︎ Setting the strategy in Databricks AutoML

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.