What you will be able to do
- Use the drop parameter to avoid co-linearity for non-regularized linear models, and explain how it interacts with unknown categories
- Predict how OneHotEncoder encodes missing values such as np.nan and None
- Limit the columns produced by high-cardinality features with min_frequency and max_categories
- Apply OneHotEncoder to selected columns inside a ColumnTransformer to avoid data leakage
1.Dropping a column to avoid co-linearity
scikit-learn's OneHotEncoder turns each categorical feature into one binary column per category. In every row exactly one of a feature's columns is 1, so those columns always add up to 1. That makes them co-linear. For some models this matters a great deal. The scikit-learn docs single out non-regularized regression (LinearRegression), where co-linearity would make the covariance matrix non-invertible.
The drop parameter removes one category per feature, so each feature produces n_categories - 1 columns. The dropped category is still recoverable: it is the row whose remaining columns are all zero. drop='first' drops the first category of every feature. If you only want this for features with exactly two categories, set drop='if_binary'. In the docs example that turns a binary gender feature into a single column while two three-category features keep all three columns each.
Checkpoint 1 of 5· Fill the gap
Which value of drop makes every feature produce n_categories - 1 columns in this sample?
>>> drop_enc = preprocessing.OneHotEncoder(drop=' ? ').fit(X)
>>> drop_enc.categories_drop='first' removes the first category of every feature. 'if_binary' applies only to two-category features, and the other two are handle_unknown values, not drop values.
Source: scikit-learn.orgdrop has a side effect when you combine it with handle_unknown='ignore'. An unseen category is encoded as all zeros, and that is exactly how the dropped category is encoded too. The two become impossible to tell apart. When inverse_transform meets an all-zero row, it returns the dropped category if one was dropped, and None if not.
>>> drop_enc = preprocessing.OneHotEncoder(drop='first',
... handle_unknown='ignore').fit(X)
>>> X_test = [['unknown', 'America', 'IE']]
>>> drop_enc.transform(X_test).toarray()
array([[0., 0., 0., 0., 0.]])Checkpoint 2 of 5· Check yourself
An encoder uses drop='first' and handle_unknown='ignore'. A test row contains a browser value never seen in training. How is it encoded?
With handle_unknown='ignore' and drop set, unknown values become all zeros. That is the dropped category's encoding, so the two are indistinguishable.
“This means that unknown categories will have the same mapping as the dropped category.”Source: scikit-learn.org
2.Missing values and high-cardinality features
Missing labels do not need a separate imputation step before one-hot encoding. OneHotEncoder handles missing values by treating them as one more category, so they get their own column. np.nan and None are not merged: if a feature contains both, each becomes its own category.
>>> X = [['Safari'], [None], [np.nan], ['Firefox']]
>>> enc = preprocessing.OneHotEncoder(handle_unknown='error').fit(X)
>>> enc.categories_
[array(['Firefox', 'Safari', None, nan], dtype=object)]High cardinality is a different problem. One column per category means a feature with many rare values produces many sparse columns. Two parameters group rare values into a single infrequent output. min_frequency sets the threshold for 'rare': an integer is a minimum count, and a float is a fraction of the total number of samples. The default of 1 encodes every category separately. max_categories caps the number of output columns per input feature, and that cap includes the combined infrequent column. In the example below, X holds 20 cats, 10 rabbits, 5 dogs and 3 snakes. With max_categories=2, only 'cat' keeps its own column and every other animal shares the infrequent column.
The sources used here do not say which model families need one-hot encoding and which can use categorical features directly. Check that question against your model's own documentation.
>>> enc = preprocessing.OneHotEncoder(max_categories=2, sparse_output=False)
>>> enc = enc.fit(X)
>>> enc.transform([['dog'], ['cat'], ['rabbit'], ['snake']])
array([[0., 1.],
[1., 0.],
[0., 1.],
[0., 1.]])Checkpoint 3 of 5· Match them up
Match each OneHotEncoder parameter to what it does
Tap a term, then the definition that fits it.
min_frequency and max_categories both control infrequent grouping. drop addresses co-linearity, and handle_unknown governs unseen values.
“This parameter sets an upper limit to the number of output features for each input feature.”Source: scikit-learn.org
Checkpoint 4 of 5· Exam question
An analyst one-hot encodes a `region` column with three levels (`north`, `south`, `east`) using `OneHotEncoder`'s default settings, producing three binary columns, and then fits a logistic regression that includes an intercept term. The model raises a warning about a singular or near-singular design matrix. What is the most likely cause and fix?
Correct answer: A — The three binary columns sum to one for every row, which is perfectly collinear with the model's intercept; refitting the encoder with `drop='first'` removes one redundant column and breaks that dependency.
- A. With no dropped category, the three indicator columns always sum to exactly one, which is a linear combination equal to the intercept column, and `drop='first'` removes that redundancy by keeping only two informative columns.
- B. `handle_unknown` only controls what happens to categories not seen during fitting; it has no effect on the dummy-variable-trap collinearity that a fully expanded one-hot encoding creates alongside an intercept.
- C. Replacing the encoding with a single ordinal integer column would remove the collinearity, but it does so by imposing a false ordering on the unordered `region` categories, which distorts the coefficients a linear model learns.
- D. Regularization strength controls how much coefficients are penalized during optimization, but it does not repair an exactly singular design matrix caused by redundant one-hot columns and an intercept.
3.Encoding inside a leakage-safe pipeline
Real tables mix types. Categorical columns need one-hot encoding while numeric or text columns need other treatment. Preprocessing everything up front, for example in pandas, is risky. Any statistics that the preprocessors learn from test data feed into the model, and that makes cross-validation scores unreliable. This is data leakage. ColumnTransformer solves it by applying a different transformer to each column inside a Pipeline that is safe from leakage and can be tuned with parameter search.
One practical detail: OneHotEncoder expects 2D input, so you name its column as a list (['city']). CountVectorizer expects 1D input and takes a plain string ('title'). Columns you do not name are dropped by default (remainder='drop'). Set remainder='passthrough' to keep them.
>>> column_trans = ColumnTransformer(
... [('categories', OneHotEncoder(dtype='int'), ['city']),
... ('title_bow', CountVectorizer(), 'title')],
... remainder='drop', verbose_feature_names_out=False)The city feature had three categories in the training data, so OneHotEncoder produced three binary columns, each named after the category it represents.
Checkpoint 5 of 5· Check yourself
In a ColumnTransformer, how should you specify the 'city' column for OneHotEncoder?
OneHotEncoder expects 2D data, so the column must be given as a list of strings. A bare string gives the 1D input that CountVectorizer expects.
“you need to specify the column as a list of strings (['city'])”Source: scikit-learn.org
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Combining drop='first' with handle_unknown='ignore' keeps unseen categories distinguishable from known ones.Why is that wrong?
Unknown values are encoded as all zeros, which is exactly how the dropped category is encoded, so the two cannot be told apart.
Covered in Dropping a column to avoid co-linearity
2.max_categories=3 gives three regular category columns plus a separate infrequent column.Why is that wrong?
The max_categories limit includes the combined infrequent column, so max_categories=3 means at most three output columns in total.
Covered in Missing values and high-cardinality features
3.It is fine to fit the encoder and other preprocessors on the full dataset before cross-validation.Why is that wrong?
When preprocessors learn statistics from test data, cross-validation scores become unreliable. Fitting inside a Pipeline with ColumnTransformer avoids this.
Covered in Encoding inside a leakage-safe pipeline
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/sparkr/overviewOfficial docs
“SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually”
↩︎ Dropping a column to avoid co-linearity - 2.https://scikit-learn.org/stable/modules/preprocessing.htmlSecondary source
“when using non-regularized regression (LinearRegression), since co-linearity would cause the covariance matrix to be non-invertible”
↩︎ Dropping a column to avoid co-linearity“you can set the parameter drop='if_binary'”
↩︎ Dropping a column to avoid co-linearity“OneHotEncoder supports categorical features with missing values by considering the missing values as an additional category”
↩︎ Missing values and high-cardinality features“If a feature contains both np.nan and None, they will be considered separate categories”
↩︎ Missing values and high-cardinality features“If min_frequency is an integer, categories with a cardinality smaller than min_frequency will be considered infrequent.”
↩︎ Missing values and high-cardinality features“This means that unknown categories will have the same mapping as the dropped category.”
↩︎ Exam trap 1“max_categories includes the feature that combines infrequent categories.”
↩︎ Exam trap 2“encode each column into n_categories - 1 columns instead of n_categories columns by using the drop parameter”
↩︎ Prediction“This means that unknown categories will have the same mapping as the dropped category.”
↩︎ Checkpoint“This parameter sets an upper limit to the number of output features for each input feature.”
↩︎ Checkpoint - 3.
“Automatic feature generation processing, like one-hot encoding for categorical features, also occurs during this stage.”
↩︎ Missing values and high-cardinality features“Validate and prepare the input table by imputing missing values and splitting data into training, validation, and test sets.”
↩︎ Encoding inside a leakage-safe pipeline - 4.https://scikit-learn.org/stable/modules/compose.htmlSecondary source
“The ColumnTransformer helps performing different transformations for different columns of the data, within a Pipeline that is safe from data leakage”
↩︎ Encoding inside a leakage-safe pipeline“Incorporating statistics from test data into the preprocessors makes cross-validation scores unreliable (known as data leakage)”
↩︎ Exam trap 3“you need to specify the column as a list of strings (['city'])”
↩︎ Checkpoint