What you will be able to do
- Explain why integer codes mislead many estimators on nominal categories and how one-hot encoding fixes that
- Fit a scikit-learn OneHotEncoder and read its output columns and categories_ attribute
- Choose between declaring categories explicitly and using handle_unknown for categories not seen in training
- Recognise when a Databricks tool performs one-hot encoding for you
Key concept
One-hot (one-of-K, dummy) encoding — A way to represent a categorical feature as numbers without inventing an order. Each of the feature's possible values gets its own binary column, and each row has a 1 in exactly one of those columns and 0 in all the others.
1.Why integer codes are not enough
Raw data often holds features as labels, not numbers: a gender, a continent, a web browser. The simplest fix is to give each label an integer code. scikit-learn's OrdinalEncoder does exactly this. It turns each categorical feature into one integer column running from 0 to n_categories - 1, so ['female', 'from US', 'uses Safari'] becomes [0., 1., 1.].
That is the problem with integer codes. The scikit-learn docs note that this representation cannot be used directly with every estimator, because many expect continuous input and will treat the codes as ordered, even though the set of browsers was ordered arbitrarily. If an estimator learns that Safari (2) is 'twice' Chrome (1), it has learned something false.
One-hot encoding, also called one-of-K or dummy encoding, removes the false order. It does not squeeze a feature into one integer column. Instead it gives every possible value its own binary column. A feature with n_categories values becomes n_categories columns, and each row has a single 1 in the column for its value. No column is 'bigger' than another, so there is no order for a model to misread.
| Aspect | OrdinalEncoder | OneHotEncoder |
|---|---|---|
| Output columns per input feature | One integer column (0 to n_categories - 1) | n_categories binary columns |
| Values in a row | A single integer code | One 1, all others 0 |
| Implied order between categories | Yes. Codes are read as ordered. | None |
| Other names | Integer codes | One-of-K, dummy encoding |
Checkpoint 1 of 4· Check yourself
A colleague encodes a 'city' feature (London, Paris, Sallisaw) as 0, 1, 2 and feeds it to an estimator that expects continuous input. What is the main risk?
Integer codes carry an implied order. For nominal labels like cities that order is meaningless, and one-hot encoding avoids it.
“would interpret the categories as being ordered, which is often not desired”Source: scikit-learn.org
2.Fitting OneHotEncoder and reading its output
In scikit-learn, one-hot encoding comes from preprocessing.OneHotEncoder. Like the other transformers, you call fit to learn the categories and then transform to produce the binary columns. The example below fits on two rows with three features: gender, location and browser. Each feature has two observed values, so the output has 2 + 2 + 2 = 6 columns.
>>> enc = preprocessing.OneHotEncoder()
>>> X = [['male', 'from US', 'uses Safari'], ['female', 'from Europe', 'uses Firefox']]
>>> enc.fit(X)
OneHotEncoder()
>>> enc.transform([['female', 'from US', 'uses Safari'],
... ['male', 'from Europe', 'uses Safari']]).toarray()
array([[1., 0., 0., 1., 0., 1.],
[0., 1., 1., 0., 0., 1.]])Columns 1–2 are gender (female, male), columns 3–4 are location (from Europe, from US) and columns 5–6 are browser (uses Firefox, uses Safari). The row is 'female', 'from US', 'uses Safari', so each pair has a 1 in exactly one position.
How do you know the column order? By default the encoder works out each feature's possible values from the training data and stores them in the categories_ attribute. Here that is ['female', 'male'], ['from Europe', 'from US'] and ['uses Firefox', 'uses Safari']. Reading categories_ is how you map output columns back to labels.
Checkpoint 2 of 4· Check yourself
After calling enc.fit(X) on a OneHotEncoder, where can you see which values each feature was found to take?
The values the encoder found during fit are stored in categories_, one array per input feature.
“By default, the values each feature can take is inferred automatically from the dataset and can be found in the categories_ attribute”Source: scikit-learn.org
You will not always write this step yourself on Databricks. The SparkR overview says that when you use SparkML GLM, SparkR one-hot encodes categorical features automatically, because under the hood SparkR trains with MLlib. Databricks AutoML forecasting also runs one-hot encoding of categorical features as part of its preprocessing stage. Knowing what the encoding does still matters, because it explains the extra columns these tools produce.
3.Categories the encoder never saw in training
The categories learned during fit are fixed from then on. Real training data can easily miss a category that appears later, and scikit-learn gives you two ways to plan for that.
The first is to list every possible value up front with the categories parameter. The docs pass four continents and four browsers, even though the training rows contain only two of each. Every declared value gets a column whether or not it appeared in training, so transforming ['female', 'from Asia', 'uses Chrome'] gives a 10-column row with a 1 in the 'from Asia' and 'uses Chrome' positions.
The second is often the better choice. When you cannot be sure the training data covers every category, the docs recommend handle_unknown='infrequent_if_exist' rather than setting the categories by hand. With this setting, an unseen value does not raise an error. Its columns for that feature are all zeros, or the value goes into an infrequent category if infrequent grouping is turned on.
>>> enc = preprocessing.OneHotEncoder(handle_unknown='infrequent_if_exist')
>>> X = [['male', 'from US', 'uses Safari'], ['female', 'from Europe', 'uses Firefox']]
>>> enc.fit(X)
OneHotEncoder(handle_unknown='infrequent_if_exist')
>>> enc.transform([['female', 'from Asia', 'uses Chrome']]).toarray()
array([[1., 0., 0., 0., 0., 0.]])Checkpoint 3 of 4· Check yourself
An encoder fitted with handle_unknown='infrequent_if_exist' (no infrequent grouping enabled) is asked to transform a browser value it never saw. What happens?
With this setting, unknown values do not raise an error. The feature's one-hot columns are simply all zeros.
“no error will be raised but the resulting one-hot encoded columns for this feature will be all zeros”Source: scikit-learn.org
Checkpoint 4 of 4· Exam question
A model is trained with an `OneHotEncoder` fit on a `device_type` column that contained `phone`, `tablet`, and `desktop` during training. In production, an inference batch arrives containing a new value, `wearable`, that the encoder never saw during fitting. The team wants the pipeline to score these rows without crashing. What should they configure on the fitted encoder?
Correct answer: D — Set `handle_unknown='ignore'` on the fitted encoder so a category absent from training is encoded as an all-zero vector at transform time instead of raising an error.
- A. The default `handle_unknown='error'` still raises on the unseen value, and catching that error only to drop the row means production requests silently go unscored instead of being handled.
- B. Refitting per batch changes the learned category order and count each time, which misaligns the encoder's output columns with the feature layout the model was actually trained on.
- C. Casting a categorical column to a numeric type does not change how `OneHotEncoder` matches values against its learned `categories_`, so a genuinely new value still fails to match and still triggers the same error.
- D. `handle_unknown='ignore'` tells the fitted encoder to represent any category outside its learned vocabulary as all zeros, letting the row pass through `transform` and reach the model instead of raising an exception.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Integer-coding a nominal feature with OrdinalEncoder is a safe replacement for one-hot encoding with any estimator.Why is that wrong?
Estimators that expect continuous input treat integer codes as ordered. For nominal labels such as browsers, that order is arbitrary and misleading.
Covered in Why integer codes are not enough
2.The only way to cope with categories missing from the training data is to list them all with the categories parameter.Why is that wrong?
The docs say handle_unknown='infrequent_if_exist' is often the better choice. It encodes unseen values as all zeros, or as an infrequent category, without raising an error.
Covered in Categories the encoder never saw in training
3.With SparkR's SparkML GLM you must one-hot encode string features yourself before fitting.Why is that wrong?
SparkR does this automatically for SparkML GLM, so you do not need to encode those features manually.
Covered in Fitting OneHotEncoder and reading its output
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.https://docs.databricks.com/aws/en/sparkr/overviewOfficial docs
“SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually”
↩︎ Why integer codes are not enough“Under the hood, SparkR uses MLlib to train the model.”
↩︎ Fitting OneHotEncoder and reading its output“SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually”
↩︎ Exam trap 3 - 2.https://scikit-learn.org/stable/modules/preprocessing.htmlSecondary source
“the set of browsers was ordered arbitrarily”
↩︎ Why integer codes are not enough“one-of-K, also known as one-hot or dummy encoding”
↩︎ Why integer codes are not enough“By default, the values each feature can take is inferred automatically from the dataset and can be found in the categories_ attribute”
↩︎ Fitting OneHotEncoder and reading its output“it can often be better to specify handle_unknown='infrequent_if_exist' instead of setting the categories manually as above”
↩︎ Categories the encoder never saw in training“transforms each categorical feature with n_categories possible values into n_categories binary features, with one of them 1, and all others 0”
↩︎ Key concept“would interpret the categories as being ordered, which is often not desired”
↩︎ Exam trap 1“it can often be better to specify handle_unknown='infrequent_if_exist' instead of setting the categories manually as above”
↩︎ Exam trap 2“would interpret the categories as being ordered, which is often not desired”
↩︎ Prediction“no error will be raised but the resulting one-hot encoded columns for this feature will be all zeros”
↩︎ Checkpoint - 3.
“Automatic feature generation processing, like one-hot encoding for categorical features, also occurs during this stage.”
↩︎ Fitting OneHotEncoder and reading its output“Automatic feature generation processing, like one-hot encoding for categorical features, also occurs during this stage.”
↩︎ Categories the encoder never saw in training