CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 2 · Lesson 26/48

    One-Hot Encoding: When to Use It and When Not To

    Identify and explain the model types or data sets for which one-hot encoding is or is not appropriate.

    14 min read
    2.08% of exam
    6 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what one-hot encoding produces and why integer category codes can mislead some models
    • Identify the model types that need one-hot encoding (linear models, GLMs, distance- and kernel-based models), and recognise when to drop a column to avoid co-linearity
    • Explain why tree-based models and label columns usually do not need one-hot encoding
    • Recognise high-cardinality features where one-hot encoding is a poor choice, and name alternatives such as target encoding or grouping infrequent categories

    Key concept

    One-hot encoding — Replacing one categorical column that has k distinct values with k binary indicator columns. Exactly one indicator is 1 in each row, so the model sees which category is present without assuming any order or distance between categories.

    1.What one-hot encoding fixes

    Most learning algorithms do arithmetic on numbers, but many useful features are categories: a browser name, a country, a product type. The simplest way to make a category numeric is to give each value an integer code, such as Firefox = 0, Chrome = 1, Safari = 2. That code is convenient, but it carries a hidden claim. To a model that treats its inputs as continuous quantities, Safari is now "twice" Chrome, and Chrome sits "between" Firefox and Safari. Those relationships come from the arbitrary order in which the codes were assigned, not from the data.

    scikit-learn's preprocessing guide states the problem directly. An integer encoding "can, however, not be used directly with all scikit-learn estimators" because those estimators expect continuous input and "would interpret the categories as being ordered, which is often not desired." One-hot encoding (also called one-of-K or dummy encoding) avoids this. Each category gets its own 0/1 column, so no category is larger than another and no false ordering exists. With scikit-learn you fit an OneHotEncoder on the categorical columns, and it learns the categories present in the data:

    Fitting scikit-learn's OneHotEncoder on three categorical features (sex, location, browser)python
    >>> enc = preprocessing.OneHotEncoder()
    >>> X = [['male', 'from US', 'uses Safari'], ['female', 'from Europe', 'uses Firefox']]
    >>> enc.fit(X)
    OneHotEncoder()

    Before you can choose an encoding, you need to know which columns are categorical in the first place. A column's data type won't always tell you. Codes such as store IDs, ZIP codes or product numbers are often stored as integers even though they are labels. Databricks AutoML checks for this: in Databricks Runtime 10.1 ML and above, its semantic type detection treats "numeric columns that contain categorical IDs" as categorical features. You can also set the semantic type yourself with a column annotation. The first question for any column is therefore whether its values are quantities or labels.

    Checkpoint 1 of 7· Check yourself

    On Databricks Runtime 10.1 ML or above, AutoML is given a regression dataset that contains an integer column of store IDs. With semantic type detection left on, how does AutoML try to treat that column?

    Sources12

    2.Models that need one-hot encoding

    One-hot encoding matters most for models that compute weighted sums of their inputs or distances between rows. Examples are linear and logistic regression, other generalized linear models (GLMs), support vector machines, k-nearest neighbours and neural networks. These models read a feature's numeric value as a magnitude. An arbitrary integer code would make them learn one slope across meaningless numbers. With one-hot columns, a linear model learns a separate coefficient for each category, which is what you want.

    Some APIs do this step for you. In the SparkR documentation on Databricks, GLMs encode categorical features automatically: "When using SparkML GLM SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually." SparkR itself is deprecated in Databricks Runtime 16.0 and above (Databricks recommends sparklyr), but the point holds generally. Linear models need categories expanded into indicator columns, whether you do it or the library does.

    There is one refinement for linear models. If you keep all k indicator columns and the model also has an intercept, the columns always add up to 1, so any one of them can be computed from the others. scikit-learn's drop parameter removes one category per feature, producing k − 1 columns. The documentation says this "is useful to avoid co-linearity in the input matrix in some classifiers" and is useful "when using non-regularized regression (LinearRegression), since co-linearity would cause the covariance matrix to be non-invertible." Regularized models are more tolerant of the redundant column. A binary category is the simplest case: it needs only one 0/1 column, and drop='if_binary' produces exactly that.

    Checkpoint 2 of 7· Fill the gap

    You will fit an unregularized LinearRegression and want the encoder to produce n_categories − 1 columns for each feature. Which parameter completes this line?

    >>> drop_enc = preprocessing.OneHotEncoder( ? ='first').fit(X)

    The Databricks getting-started tutorial uses this one-column pattern for a two-valued category. The wine dataset comes in red and white files, and instead of keeping a color string, the notebook adds one indicator column:

    A two-valued category (red vs. white wine) encoded as a single 0/1 column in the Databricks ML getting-started tutorialpython
    # Add Boolean fields for red and white wine
    white_wine['is_red'] = 0.0
    red_wine['is_red'] = 1.0
    data_df = pd.concat([white_wine, red_wine], axis=0)

    Checkpoint 3 of 7· Exam question

    A data scientist at a retail bank is training a scikit-learn `LogisticRegression` model to predict customer churn using Databricks. One input feature, `contract_type`, holds three unordered categories: month-to-month, one-year, and two-year. Which preprocessing approach best prepares this feature for the model?

    Sources32

    3.When one-hot encoding is unnecessary: trees and labels

    Tree-based models work differently from linear models. Decision trees, random forests and gradient-boosted trees (such as XGBoost) don't multiply features by coefficients. They split rows on thresholds. With integer codes, two splits ("code ≤ 0.5", then "code ≤ 1.5") are enough to isolate any single category, so an arbitrary order does far less harm. Some tree implementations, including Spark MLlib's tree learners, can also use indexed categorical features directly. For tree ensembles, an ordinal or index encoding is therefore usually enough.

    One-hot encoding can even make trees worse. A category with many values turns into many sparse 0/1 columns. Each split can then only separate one category from all the rest, the trees grow deeper, and the feature's importance is spread across dozens of weak columns. That's why "tree-based models don't need one-hot encoding" is a standard rule of thumb. Databricks' own MLlib examples include gradient-boosted trees used for "regression using gradient boosted trees to predict bike rental counts (per hour) from information such as day of the week, weather, season". Those inputs are categorical by nature, and they are the kind of data trees handle without expansion.

    The other place one-hot encoding is usually unnecessary is the label (the target column). Classifiers generally expect classes as integer IDs, not as indicator vectors. The Databricks guide to fine-tuning Hugging Face models makes this concrete: AutoModelForSequenceClassification is a model loader "which expects integer IDs as the category labels". The guide builds label2id and id2label mappings from the string labels instead of expanding them into columns.

    Checkpoint 4 of 7· Check yourself

    You are preparing a Spark DataFrame of text and string labels ("positive", "negative", "neutral") to fine-tune a Hugging Face AutoModelForSequenceClassification model on Databricks. How should the label column be represented?

    Checkpoint 5 of 7· Exam question

    An ML engineer is training a `sklearn.ensemble.GradientBoostingClassifier` on a dataset where the `region` feature has five unordered categories. A teammate insists every categorical feature must be one-hot encoded before training any model. How should the engineer respond?

    Sources45

    4.When the data rules it out: high-cardinality features

    Even for a linear model, the data itself can make one-hot encoding a poor choice. Cardinality is the number of distinct values a column has. Encoding a 30,000-value ZIP code column adds 30,000 columns, and in every row 29,999 of them are 0. The matrix becomes very wide and very sparse. Rare categories appear in only a few rows, so their coefficients are poorly estimated. Categories that never appeared in training have no column at all.

    scikit-learn's guide names this case when it introduces TargetEncoder. Target encoding "is useful with categorical features with high cardinality, where one-hot encoding would inflate the feature space", and "A classical example of high cardinality categories are location based such as zip code or region." A target encoder replaces each category with a single number based on the mean of the target for that category, so the column stays one column wide. Because that number comes from the label, scikit-learn's fit_transform uses cross fitting to prevent target information from leaking into the training representation.

    If you want to keep one-hot encoding, you can reduce the cardinality first. OneHotEncoder can group rare values: min_frequency sends categories below a frequency threshold into a single "infrequent" column, and max_categories caps the total number of output columns. Other options include an OrdinalEncoder (for tree models) and learned embeddings. If you do keep a wide one-hot output, it is sparse, and Databricks Feature Store notes that "You can store sparse vectors, tensors, and embeddings as MapType."

    Encoding choices in scikit-learn and the situations each one suits
    Encoder / optionOutputGood fitPoor fit
    OneHotEncodern_categories binary columns per featureLow-cardinality nominal features for linear and distance-based modelsHigh-cardinality features such as zip code. Often unnecessary for trees.
    OneHotEncoder with dropn_categories − 1 columns per featureNon-regularized LinearRegression, where co-linearity breaks the fitCases where an unknown category must stay distinct from the dropped one
    OneHotEncoder with min_frequency / max_categoriesRare categories grouped into one infrequent columnModerate cardinality with a long tail of rare valuesFeatures where the rare categories carry the signal
    OrdinalEncoderOne integer column (0 to n_categories − 1)Tree-based models. Genuinely ordered categories.Linear models on nominal categories, which would read the codes as ordered
    TargetEncoderOne column of category-conditioned target meansHigh-cardinality nominal featuresSmall data with many rare categories if cross fitting is skipped (risk of leakage)

    Checkpoint 6 of 7· Match them up

    Match each scikit-learn encoder or option to what it does

    Tap a term, then the definition that fits it.

    Checkpoint 7 of 7· Exam question

    A team is building a `RandomForestRegressor` to predict delivery times using a `delivery_zip_code` feature containing roughly 15,000 unique ZIP codes. Applying `OneHotEncoder` to this feature during preprocessing produces a dataset that fails to fit in the cluster's memory. What is the most appropriate fix?

    Sources62

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A column stored as integers is numeric, so it can go straight into the model as a continuous feature.Why is that wrong?

      Integer columns often hold identifiers such as store IDs or ZIP codes. These are categories and need categorical handling. Databricks AutoML detects exactly this case.

      Covered in What one-hot encoding fixes

    2. 2.Integer category codes are wrong for every model, so every categorical feature must be one-hot encoded.Why is that wrong?

      Integer codes only mislead estimators that treat their inputs as continuous quantities. Tree-based models split on thresholds and usually handle ordinal or indexed codes well.

      Covered in What one-hot encoding fixes

    3. 3.When one-hot encoding for plain (unregularized) linear regression, keep all k indicator columns per feature.Why is that wrong?

      With an intercept, all k columns are perfectly collinear. Drop one category per feature, for example with drop='first'.

      Covered in Models that need one-hot encoding

    4. 4.One-hot encoding is the safe default for any categorical feature, including zip code or region.Why is that wrong?

      For high-cardinality features, one-hot encoding inflates the feature space. Target encoding, or grouping infrequent categories, is usually the better choice.

      Covered in When the data rules it out: high-cardinality features

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Numeric columns that contain categorical IDs are treated as a categorical feature.”
      ↩︎ What one-hot encoding fixes
      “Numeric columns that contain categorical IDs are treated as a categorical feature.”
      ↩︎ Exam trap 1
    2. 2.
      “would interpret the categories as being ordered, which is often not desired”
      ↩︎ What one-hot encoding fixes
      “This is useful to avoid co-linearity in the input matrix in some classifiers.”
      ↩︎ Models that need one-hot encoding
      “A classical example of high cardinality categories are location based such as zip code or region.”
      ↩︎ When the data rules it out: high-cardinality features
      “Similarly, OneHotEncoder can be configured to group together infrequent categories:”
      ↩︎ When the data rules it out: high-cardinality features
      “transforms each categorical feature with n_categories possible values into n_categories binary features, with one of them 1, and all others 0.”
      ↩︎ Key concept
      “not be used directly with all scikit-learn estimators, as these expect continuous input”
      ↩︎ Exam trap 2
      “since co-linearity would cause the covariance matrix to be non-invertible”
      ↩︎ Exam trap 3
      “This encoding scheme is useful with categorical features with high cardinality, where one-hot encoding would inflate the feature space”
      ↩︎ Exam trap 4
      “The TargetEncoder uses the target mean conditioned on the categorical feature for encoding unordered categories”
      ↩︎ Checkpoint
    3. 3.
      “When using SparkML GLM SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually.”
      ↩︎ Models that need one-hot encoding
      “SparkR in Databricks is deprecated in Databricks Runtime 16.0 and above.”
      ↩︎ Models that need one-hot encoding
    4. 4.
      “regression using gradient boosted trees to predict bike rental counts (per hour) from information such as day of the week, weather, season”
      ↩︎ When one-hot encoding is unnecessary: trees and labels
    5. 6.
      “You can store sparse vectors, tensors, and embeddings as MapType.”
      ↩︎ When the data rules it out: high-cardinality features

    Spotted a mistake, or was something unclear? Tell us.