CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 2 · Lesson 25/48

    Tuning OneHotEncoder: drop, Missing Values, Infrequent Categories and Pipelines

    Use one-hot encoding for categorical features

    8 min read
    2.08% of exam
    4 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Use the drop parameter to avoid co-linearity for non-regularized linear models, and explain how it interacts with unknown categories
    • Predict how OneHotEncoder encodes missing values such as np.nan and None
    • Limit the columns produced by high-cardinality features with min_frequency and max_categories
    • Apply OneHotEncoder to selected columns inside a ColumnTransformer to avoid data leakage

    1.Dropping a column to avoid co-linearity

    scikit-learn's OneHotEncoder turns each categorical feature into one binary column per category. In every row exactly one of a feature's columns is 1, so those columns always add up to 1. That makes them co-linear. For some models this matters a great deal. The scikit-learn docs single out non-regularized regression (LinearRegression), where co-linearity would make the covariance matrix non-invertible.

    The drop parameter removes one category per feature, so each feature produces n_categories - 1 columns. The dropped category is still recoverable: it is the row whose remaining columns are all zero. drop='first' drops the first category of every feature. If you only want this for features with exactly two categories, set drop='if_binary'. In the docs example that turns a binary gender feature into a single column while two three-category features keep all three columns each.

    Checkpoint 1 of 5· Fill the gap

    Which value of drop makes every feature produce n_categories - 1 columns in this sample?

    >>> drop_enc = preprocessing.OneHotEncoder(drop=' ? ').fit(X)
    >>> drop_enc.categories_

    drop has a side effect when you combine it with handle_unknown='ignore'. An unseen category is encoded as all zeros, and that is exactly how the dropped category is encoded too. The two become impossible to tell apart. When inverse_transform meets an all-zero row, it returns the dropped category if one was dropped, and None if not.

    drop='first' with handle_unknown='ignore': every unknown value becomes all zeros, the same encoding as the dropped categorypython
    >>> drop_enc = preprocessing.OneHotEncoder(drop='first',
    ...                                        handle_unknown='ignore').fit(X)
    >>> X_test = [['unknown', 'America', 'IE']]
    >>> drop_enc.transform(X_test).toarray()
    array([[0., 0., 0., 0., 0.]])

    Checkpoint 2 of 5· Check yourself

    An encoder uses drop='first' and handle_unknown='ignore'. A test row contains a browser value never seen in training. How is it encoded?

    Sources12

    2.Missing values and high-cardinality features

    Missing labels do not need a separate imputation step before one-hot encoding. OneHotEncoder handles missing values by treating them as one more category, so they get their own column. np.nan and None are not merged: if a feature contains both, each becomes its own category.

    None and np.nan become two separate categories, each with its own output columnpython
    >>> X = [['Safari'], [None], [np.nan], ['Firefox']]
    >>> enc = preprocessing.OneHotEncoder(handle_unknown='error').fit(X)
    >>> enc.categories_
    [array(['Firefox', 'Safari', None, nan], dtype=object)]

    High cardinality is a different problem. One column per category means a feature with many rare values produces many sparse columns. Two parameters group rare values into a single infrequent output. min_frequency sets the threshold for 'rare': an integer is a minimum count, and a float is a fraction of the total number of samples. The default of 1 encodes every category separately. max_categories caps the number of output columns per input feature, and that cap includes the combined infrequent column. In the example below, X holds 20 cats, 10 rabbits, 5 dogs and 3 snakes. With max_categories=2, only 'cat' keeps its own column and every other animal shares the infrequent column.

    The sources used here do not say which model families need one-hot encoding and which can use categorical features directly. Check that question against your model's own documentation.

    max_categories=2: 'cat' keeps its own column and dog, rabbit and snake share the infrequent columnpython
    >>> enc = preprocessing.OneHotEncoder(max_categories=2, sparse_output=False)
    >>> enc = enc.fit(X)
    >>> enc.transform([['dog'], ['cat'], ['rabbit'], ['snake']])
    array([[0., 1.],
           [1., 0.],
           [0., 1.],
           [0., 1.]])

    Checkpoint 3 of 5· Match them up

    Match each OneHotEncoder parameter to what it does

    Tap a term, then the definition that fits it.

    Checkpoint 4 of 5· Exam question

    An analyst one-hot encodes a `region` column with three levels (`north`, `south`, `east`) using `OneHotEncoder`'s default settings, producing three binary columns, and then fits a logistic regression that includes an intercept term. The model raises a warning about a singular or near-singular design matrix. What is the most likely cause and fix?

    Sources32

    3.Encoding inside a leakage-safe pipeline

    Real tables mix types. Categorical columns need one-hot encoding while numeric or text columns need other treatment. Preprocessing everything up front, for example in pandas, is risky. Any statistics that the preprocessors learn from test data feed into the model, and that makes cross-validation scores unreliable. This is data leakage. ColumnTransformer solves it by applying a different transformer to each column inside a Pipeline that is safe from leakage and can be tuned with parameter search.

    One practical detail: OneHotEncoder expects 2D input, so you name its column as a list (['city']). CountVectorizer expects 1D input and takes a plain string ('title'). Columns you do not name are dropped by default (remainder='drop'). Set remainder='passthrough' to keep them.

    ColumnTransformer: one-hot encode 'city' (passed as a list) and vectorize 'title' (passed as a string)python
    >>> column_trans = ColumnTransformer(
    ...     [('categories', OneHotEncoder(dtype='int'), ['city']),
    ...      ('title_bow', CountVectorizer(), 'title')],
    ...     remainder='drop', verbose_feature_names_out=False)

    Checkpoint 5 of 5· Check yourself

    In a ColumnTransformer, how should you specify the 'city' column for OneHotEncoder?

    Sources34

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Combining drop='first' with handle_unknown='ignore' keeps unseen categories distinguishable from known ones.Why is that wrong?

      Unknown values are encoded as all zeros, which is exactly how the dropped category is encoded, so the two cannot be told apart.

      Covered in Dropping a column to avoid co-linearity

    2. 2.max_categories=3 gives three regular category columns plus a separate infrequent column.Why is that wrong?

      The max_categories limit includes the combined infrequent column, so max_categories=3 means at most three output columns in total.

      Covered in Missing values and high-cardinality features

    3. 3.It is fine to fit the encoder and other preprocessors on the full dataset before cross-validation.Why is that wrong?

      When preprocessors learn statistics from test data, cross-validation scores become unreliable. Fitting inside a Pipeline with ColumnTransformer avoids this.

      Covered in Encoding inside a leakage-safe pipeline

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually”
      ↩︎ Dropping a column to avoid co-linearity
    2. 2.
      “when using non-regularized regression (LinearRegression), since co-linearity would cause the covariance matrix to be non-invertible”
      ↩︎ Dropping a column to avoid co-linearity
      “you can set the parameter drop='if_binary'”
      ↩︎ Dropping a column to avoid co-linearity
      “OneHotEncoder supports categorical features with missing values by considering the missing values as an additional category”
      ↩︎ Missing values and high-cardinality features
      “If a feature contains both np.nan and None, they will be considered separate categories”
      ↩︎ Missing values and high-cardinality features
      “If min_frequency is an integer, categories with a cardinality smaller than min_frequency will be considered infrequent.”
      ↩︎ Missing values and high-cardinality features
      “This means that unknown categories will have the same mapping as the dropped category.”
      ↩︎ Exam trap 1
      “max_categories includes the feature that combines infrequent categories.”
      ↩︎ Exam trap 2
      “encode each column into n_categories - 1 columns instead of n_categories columns by using the drop parameter”
      ↩︎ Prediction
      “This means that unknown categories will have the same mapping as the dropped category.”
      ↩︎ Checkpoint
      “This parameter sets an upper limit to the number of output features for each input feature.”
      ↩︎ Checkpoint
    3. 3.
      “Automatic feature generation processing, like one-hot encoding for categorical features, also occurs during this stage.”
      ↩︎ Missing values and high-cardinality features
      “Validate and prepare the input table by imputing missing values and splitting data into training, validation, and test sets.”
      ↩︎ Encoding inside a leakage-safe pipeline
    4. 4.
      “The ColumnTransformer helps performing different transformations for different columns of the data, within a Pipeline that is safe from data leakage”
      ↩︎ Encoding inside a leakage-safe pipeline
      “Incorporating statistics from test data into the preprocessors makes cross-validation scores unreliable (known as data leakage)”
      ↩︎ Exam trap 3
      “you need to specify the column as a list of strings (['city'])”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.