What you will be able to do
- Explain what one-hot encoding produces and why integer category codes can mislead some models
- Identify the model types that need one-hot encoding (linear models, GLMs, distance- and kernel-based models), and recognise when to drop a column to avoid co-linearity
- Explain why tree-based models and label columns usually do not need one-hot encoding
- Recognise high-cardinality features where one-hot encoding is a poor choice, and name alternatives such as target encoding or grouping infrequent categories
Key concept
One-hot encoding — Replacing one categorical column that has k distinct values with k binary indicator columns. Exactly one indicator is 1 in each row, so the model sees which category is present without assuming any order or distance between categories.
1.What one-hot encoding fixes
Most learning algorithms do arithmetic on numbers, but many useful features are categories: a browser name, a country, a product type. The simplest way to make a category numeric is to give each value an integer code, such as Firefox = 0, Chrome = 1, Safari = 2. That code is convenient, but it carries a hidden claim. To a model that treats its inputs as continuous quantities, Safari is now "twice" Chrome, and Chrome sits "between" Firefox and Safari. Those relationships come from the arbitrary order in which the codes were assigned, not from the data.
scikit-learn's preprocessing guide states the problem directly. An integer encoding "can, however, not be used directly with all scikit-learn estimators" because those estimators expect continuous input and "would interpret the categories as being ordered, which is often not desired." One-hot encoding (also called one-of-K or dummy encoding) avoids this. Each category gets its own 0/1 column, so no category is larger than another and no false ordering exists. With scikit-learn you fit an OneHotEncoder on the categorical columns, and it learns the categories present in the data:
>>> enc = preprocessing.OneHotEncoder()
>>> X = [['male', 'from US', 'uses Safari'], ['female', 'from Europe', 'uses Firefox']]
>>> enc.fit(X)
OneHotEncoder()Before you can choose an encoding, you need to know which columns are categorical in the first place. A column's data type won't always tell you. Codes such as store IDs, ZIP codes or product numbers are often stored as integers even though they are labels. Databricks AutoML checks for this: in Databricks Runtime 10.1 ML and above, its semantic type detection treats "numeric columns that contain categorical IDs" as categorical features. You can also set the semantic type yourself with a column annotation. The first question for any column is therefore whether its values are quantities or labels.
Checkpoint 1 of 7· Check yourself
On Databricks Runtime 10.1 ML or above, AutoML is given a regression dataset that contains an integer column of store IDs. With semantic type detection left on, how does AutoML try to treat that column?
AutoML's semantic type detection looks past the storage type. When a numeric column holds categorical IDs, AutoML treats it as categorical.
“Numeric columns that contain categorical IDs are treated as a categorical feature.”Source: docs.databricks.com
2.Models that need one-hot encoding
One-hot encoding matters most for models that compute weighted sums of their inputs or distances between rows. Examples are linear and logistic regression, other generalized linear models (GLMs), support vector machines, k-nearest neighbours and neural networks. These models read a feature's numeric value as a magnitude. An arbitrary integer code would make them learn one slope across meaningless numbers. With one-hot columns, a linear model learns a separate coefficient for each category, which is what you want.
Some APIs do this step for you. In the SparkR documentation on Databricks, GLMs encode categorical features automatically: "When using SparkML GLM SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually." SparkR itself is deprecated in Databricks Runtime 16.0 and above (Databricks recommends sparklyr), but the point holds generally. Linear models need categories expanded into indicator columns, whether you do it or the library does.
There is one refinement for linear models. If you keep all k indicator columns and the model also has an intercept, the columns always add up to 1, so any one of them can be computed from the others. scikit-learn's drop parameter removes one category per feature, producing k − 1 columns. The documentation says this "is useful to avoid co-linearity in the input matrix in some classifiers" and is useful "when using non-regularized regression (LinearRegression), since co-linearity would cause the covariance matrix to be non-invertible." Regularized models are more tolerant of the redundant column. A binary category is the simplest case: it needs only one 0/1 column, and drop='if_binary' produces exactly that.
Checkpoint 2 of 7· Fill the gap
You will fit an unregularized LinearRegression and want the encoder to produce n_categories − 1 columns for each feature. Which parameter completes this line?
>>> drop_enc = preprocessing.OneHotEncoder( ? ='first').fit(X)drop tells OneHotEncoder which category to drop from each feature. drop='first' removes the first one, which avoids co-linearity in non-regularized regression.
The Databricks getting-started tutorial uses this one-column pattern for a two-valued category. The wine dataset comes in red and white files, and instead of keeping a color string, the notebook adds one indicator column:
# Add Boolean fields for red and white wine
white_wine['is_red'] = 0.0
red_wine['is_red'] = 1.0
data_df = pd.concat([white_wine, red_wine], axis=0)Checkpoint 3 of 7· Exam question
A data scientist at a retail bank is training a scikit-learn `LogisticRegression` model to predict customer churn using Databricks. One input feature, `contract_type`, holds three unordered categories: month-to-month, one-year, and two-year. Which preprocessing approach best prepares this feature for the model?
Correct answer: A — One-hot encode `contract_type` into three binary indicator columns, since logistic regression multiplies each feature by a learned weight and would treat integer codes as ranked magnitudes.
- A. Correct — logistic regression is a linear model that computes weighted sums of feature values, so an unordered categorical feature must become binary indicators; otherwise arbitrary integer codes like 0, 1, 2 would imply a false ordinal relationship between contract lengths.
- B. Incorrect — scikit-learn's `LogisticRegression` does not detect or compensate for unordered categories; ordinal integers are treated as ordinary numeric input, so the two-year contract would be modeled as inherently three times the one-year contract.
- C. Incorrect — scikit-learn estimators require numeric input and raise an error on raw string columns; Spark's `LogisticRegression` (from `pyspark.ml`) also does not silently one-hot encode string columns without an explicit `StringIndexer`/`OneHotEncoder` stage.
- D. Incorrect — dropping the feature discards real predictive signal; contract length is a well-known churn driver, and having only three categories is exactly the low-cardinality case where one-hot encoding is cheap and appropriate rather than a reason to remove the column.
3.When one-hot encoding is unnecessary: trees and labels
Tree-based models work differently from linear models. Decision trees, random forests and gradient-boosted trees (such as XGBoost) don't multiply features by coefficients. They split rows on thresholds. With integer codes, two splits ("code ≤ 0.5", then "code ≤ 1.5") are enough to isolate any single category, so an arbitrary order does far less harm. Some tree implementations, including Spark MLlib's tree learners, can also use indexed categorical features directly. For tree ensembles, an ordinal or index encoding is therefore usually enough.
One-hot encoding can even make trees worse. A category with many values turns into many sparse 0/1 columns. Each split can then only separate one category from all the rest, the trees grow deeper, and the feature's importance is spread across dozens of weak columns. That's why "tree-based models don't need one-hot encoding" is a standard rule of thumb. Databricks' own MLlib examples include gradient-boosted trees used for "regression using gradient boosted trees to predict bike rental counts (per hour) from information such as day of the week, weather, season". Those inputs are categorical by nature, and they are the kind of data trees handle without expansion.
The other place one-hot encoding is usually unnecessary is the label (the target column). Classifiers generally expect classes as integer IDs, not as indicator vectors. The Databricks guide to fine-tuning Hugging Face models makes this concrete: AutoModelForSequenceClassification is a model loader "which expects integer IDs as the category labels". The guide builds label2id and id2label mappings from the string labels instead of expanding them into columns.
Checkpoint 4 of 7· Check yourself
You are preparing a Spark DataFrame of text and string labels ("positive", "negative", "neutral") to fine-tune a Hugging Face AutoModelForSequenceClassification model on Databricks. How should the label column be represented?
The model loader expects integer class IDs. The Databricks guide builds label2id / id2label mappings and replaces the string labels with integers.
“which expects integer IDs as the category labels”Source: docs.databricks.com
Checkpoint 5 of 7· Exam question
An ML engineer is training a `sklearn.ensemble.GradientBoostingClassifier` on a dataset where the `region` feature has five unordered categories. A teammate insists every categorical feature must be one-hot encoded before training any model. How should the engineer respond?
Correct answer: A — Push back and use ordinal or native category encoding instead, because tree-based models split on individual feature thresholds and do not assume a linear relationship between encoded values.
- A. Correct — decision trees and gradient boosting choose split points per feature and evaluate categories (or ordinal codes) independently, so they do not suffer from the false-ordering problem that motivates one-hot encoding for linear models; native or ordinal encoding is usually sufficient and keeps the feature space compact.
- B. Incorrect — gradient boosting does not compute a single weighted linear combination of raw feature values the way logistic or linear regression does; each tree instead partitions the feature space with threshold splits, so one-hot encoding is not required to avoid a false numeric ordering.
- C. Incorrect — decision trees are not restricted to binary categorical inputs; a single node can effectively partition many categories through repeated splits, so hashing `region` down to two buckets discards information for no algorithmic reason.
- D. Incorrect — `max_depth` and similar hyperparameters control tree structure and regularization, not how categorical values are represented; the model still needs `region` encoded as some numeric form (ordinal or one-hot) before it can be dropped into training data at all.
4.When the data rules it out: high-cardinality features
Even for a linear model, the data itself can make one-hot encoding a poor choice. Cardinality is the number of distinct values a column has. Encoding a 30,000-value ZIP code column adds 30,000 columns, and in every row 29,999 of them are 0. The matrix becomes very wide and very sparse. Rare categories appear in only a few rows, so their coefficients are poorly estimated. Categories that never appeared in training have no column at all.
scikit-learn's guide names this case when it introduces TargetEncoder. Target encoding "is useful with categorical features with high cardinality, where one-hot encoding would inflate the feature space", and "A classical example of high cardinality categories are location based such as zip code or region." A target encoder replaces each category with a single number based on the mean of the target for that category, so the column stays one column wide. Because that number comes from the label, scikit-learn's fit_transform uses cross fitting to prevent target information from leaking into the training representation.
If you want to keep one-hot encoding, you can reduce the cardinality first. OneHotEncoder can group rare values: min_frequency sends categories below a frequency threshold into a single "infrequent" column, and max_categories caps the total number of output columns. Other options include an OrdinalEncoder (for tree models) and learned embeddings. If you do keep a wide one-hot output, it is sparse, and Databricks Feature Store notes that "You can store sparse vectors, tensors, and embeddings as MapType."
| Encoder / option | Output | Good fit | Poor fit |
|---|---|---|---|
| OneHotEncoder | n_categories binary columns per feature | Low-cardinality nominal features for linear and distance-based models | High-cardinality features such as zip code. Often unnecessary for trees. |
| OneHotEncoder with drop | n_categories − 1 columns per feature | Non-regularized LinearRegression, where co-linearity breaks the fit | Cases where an unknown category must stay distinct from the dropped one |
| OneHotEncoder with min_frequency / max_categories | Rare categories grouped into one infrequent column | Moderate cardinality with a long tail of rare values | Features where the rare categories carry the signal |
| OrdinalEncoder | One integer column (0 to n_categories − 1) | Tree-based models. Genuinely ordered categories. | Linear models on nominal categories, which would read the codes as ordered |
| TargetEncoder | One column of category-conditioned target means | High-cardinality nominal features | Small data with many rare categories if cross fitting is skipped (risk of leakage) |
Checkpoint 6 of 7· Match them up
Match each scikit-learn encoder or option to what it does
Tap a term, then the definition that fits it.
One-hot creates indicator columns, ordinal creates integer codes, and target encoding uses per-category target means. min_frequency is the OneHotEncoder option that groups rare categories.
“The TargetEncoder uses the target mean conditioned on the categorical feature for encoding unordered categories”Source: scikit-learn.org
Checkpoint 7 of 7· Exam question
A team is building a `RandomForestRegressor` to predict delivery times using a `delivery_zip_code` feature containing roughly 15,000 unique ZIP codes. Applying `OneHotEncoder` to this feature during preprocessing produces a dataset that fails to fit in the cluster's memory. What is the most appropriate fix?
Correct answer: A — Replace one-hot encoding with target encoding or another low-dimensional representation, since ZIP code is the classic high-cardinality case where one-hot encoding inflates the feature space far beyond what is useful.
- A. Correct — scikit-learn's own guidance recommends target encoding for high-cardinality categorical features like ZIP code or region, because one-hot encoding a feature with tens of thousands of categories creates that many additional sparse columns, which is expensive for any downstream model to process.
- B. Incorrect — setting `sparse_output=False` forces the encoder to materialize a dense array, which uses more memory than the default sparse CSR representation, not less; it does nothing to reduce the underlying 15,000-column dimensionality problem.
- C. Incorrect — 15,000 one-hot columns for a single feature is an unusually large and wasteful representation, not a normal ensemble training scenario; adding more driver memory treats a symptom rather than the root cause of an inflated feature space.
- D. Incorrect — reducing `n_estimators` to a single tree lowers ensemble memory use but does not shrink the width of the training matrix at all; the 15,000-column encoding still has to be constructed and held in memory before any tree can be fit.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A column stored as integers is numeric, so it can go straight into the model as a continuous feature.Why is that wrong?
Integer columns often hold identifiers such as store IDs or ZIP codes. These are categories and need categorical handling. Databricks AutoML detects exactly this case.
Covered in What one-hot encoding fixes
2.Integer category codes are wrong for every model, so every categorical feature must be one-hot encoded.Why is that wrong?
Integer codes only mislead estimators that treat their inputs as continuous quantities. Tree-based models split on thresholds and usually handle ordinal or indexed codes well.
Covered in What one-hot encoding fixes
3.When one-hot encoding for plain (unregularized) linear regression, keep all k indicator columns per feature.Why is that wrong?
With an intercept, all k columns are perfectly collinear. Drop one category per feature, for example with drop='first'.
Covered in Models that need one-hot encoding
4.One-hot encoding is the safe default for any categorical feature, including zip code or region.Why is that wrong?
For high-cardinality features, one-hot encoding inflates the feature space. Target encoding, or grouping infrequent categories, is usually the better choice.
Covered in When the data rules it out: high-cardinality features
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Numeric columns that contain categorical IDs are treated as a categorical feature.”
↩︎ What one-hot encoding fixes“Numeric columns that contain categorical IDs are treated as a categorical feature.”
↩︎ Exam trap 1 - 2.https://scikit-learn.org/stable/modules/preprocessing.htmlSecondary source
“would interpret the categories as being ordered, which is often not desired”
↩︎ What one-hot encoding fixes“This is useful to avoid co-linearity in the input matrix in some classifiers.”
↩︎ Models that need one-hot encoding“A classical example of high cardinality categories are location based such as zip code or region.”
↩︎ When the data rules it out: high-cardinality features“Similarly, OneHotEncoder can be configured to group together infrequent categories:”
↩︎ When the data rules it out: high-cardinality features“transforms each categorical feature with n_categories possible values into n_categories binary features, with one of them 1, and all others 0.”
↩︎ Key concept“not be used directly with all scikit-learn estimators, as these expect continuous input”
↩︎ Exam trap 2“since co-linearity would cause the covariance matrix to be non-invertible”
↩︎ Exam trap 3“This encoding scheme is useful with categorical features with high cardinality, where one-hot encoding would inflate the feature space”
↩︎ Exam trap 4“The TargetEncoder uses the target mean conditioned on the categorical feature for encoding unordered categories”
↩︎ Checkpoint - 3.https://docs.databricks.com/aws/en/sparkr/overviewOfficial docs
“When using SparkML GLM SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually.”
↩︎ Models that need one-hot encoding“SparkR in Databricks is deprecated in Databricks Runtime 16.0 and above.”
↩︎ Models that need one-hot encoding - 4.
“regression using gradient boosted trees to predict bike rental counts (per hour) from information such as day of the week, weather, season”
↩︎ When one-hot encoding is unnecessary: trees and labels - 5.
“which expects integer IDs as the category labels”
↩︎ When one-hot encoding is unnecessary: trees and labels - 6.
“You can store sparse vectors, tensors, and embeddings as MapType.”
↩︎ When the data rules it out: high-cardinality features