CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 2 · Lesson 25/48

    One-Hot Encoding Categorical Features with OneHotEncoder

    Use one-hot encoding for categorical features

    9 min read
    2.08% of exam
    3 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why integer codes mislead many estimators on nominal categories and how one-hot encoding fixes that
    • Fit a scikit-learn OneHotEncoder and read its output columns and categories_ attribute
    • Choose between declaring categories explicitly and using handle_unknown for categories not seen in training
    • Recognise when a Databricks tool performs one-hot encoding for you

    Key concept

    One-hot (one-of-K, dummy) encoding — A way to represent a categorical feature as numbers without inventing an order. Each of the feature's possible values gets its own binary column, and each row has a 1 in exactly one of those columns and 0 in all the others.

    1.Why integer codes are not enough

    Raw data often holds features as labels, not numbers: a gender, a continent, a web browser. The simplest fix is to give each label an integer code. scikit-learn's OrdinalEncoder does exactly this. It turns each categorical feature into one integer column running from 0 to n_categories - 1, so ['female', 'from US', 'uses Safari'] becomes [0., 1., 1.].

    That is the problem with integer codes. The scikit-learn docs note that this representation cannot be used directly with every estimator, because many expect continuous input and will treat the codes as ordered, even though the set of browsers was ordered arbitrarily. If an estimator learns that Safari (2) is 'twice' Chrome (1), it has learned something false.

    One-hot encoding, also called one-of-K or dummy encoding, removes the false order. It does not squeeze a feature into one integer column. Instead it gives every possible value its own binary column. A feature with n_categories values becomes n_categories columns, and each row has a single 1 in the column for its value. No column is 'bigger' than another, so there is no order for a model to misread.

    OrdinalEncoder and OneHotEncoder side by side, as the scikit-learn preprocessing docs describe them
    AspectOrdinalEncoderOneHotEncoder
    Output columns per input featureOne integer column (0 to n_categories - 1)n_categories binary columns
    Values in a rowA single integer codeOne 1, all others 0
    Implied order between categoriesYes. Codes are read as ordered.None
    Other namesInteger codesOne-of-K, dummy encoding

    Checkpoint 1 of 4· Check yourself

    A colleague encodes a 'city' feature (London, Paris, Sallisaw) as 0, 1, 2 and feeds it to an estimator that expects continuous input. What is the main risk?

    Sources12

    2.Fitting OneHotEncoder and reading its output

    In scikit-learn, one-hot encoding comes from preprocessing.OneHotEncoder. Like the other transformers, you call fit to learn the categories and then transform to produce the binary columns. The example below fits on two rows with three features: gender, location and browser. Each feature has two observed values, so the output has 2 + 2 + 2 = 6 columns.

    Fit OneHotEncoder on three categorical features, then transform two rows into six binary columnspython
    >>> enc = preprocessing.OneHotEncoder()
    >>> X = [['male', 'from US', 'uses Safari'], ['female', 'from Europe', 'uses Firefox']]
    >>> enc.fit(X)
    OneHotEncoder()
    >>> enc.transform([['female', 'from US', 'uses Safari'],
    ...                ['male', 'from Europe', 'uses Safari']]).toarray()
    array([[1., 0., 0., 1., 0., 1.],
           [0., 1., 1., 0., 0., 1.]])

    How do you know the column order? By default the encoder works out each feature's possible values from the training data and stores them in the categories_ attribute. Here that is ['female', 'male'], ['from Europe', 'from US'] and ['uses Firefox', 'uses Safari']. Reading categories_ is how you map output columns back to labels.

    Checkpoint 2 of 4· Check yourself

    After calling enc.fit(X) on a OneHotEncoder, where can you see which values each feature was found to take?

    You will not always write this step yourself on Databricks. The SparkR overview says that when you use SparkML GLM, SparkR one-hot encodes categorical features automatically, because under the hood SparkR trains with MLlib. Databricks AutoML forecasting also runs one-hot encoding of categorical features as part of its preprocessing stage. Knowing what the encoding does still matters, because it explains the extra columns these tools produce.

    Sources312

    3.Categories the encoder never saw in training

    The categories learned during fit are fixed from then on. Real training data can easily miss a category that appears later, and scikit-learn gives you two ways to plan for that.

    The first is to list every possible value up front with the categories parameter. The docs pass four continents and four browsers, even though the training rows contain only two of each. Every declared value gets a column whether or not it appeared in training, so transforming ['female', 'from Asia', 'uses Chrome'] gives a 10-column row with a 1 in the 'from Asia' and 'uses Chrome' positions.

    The second is often the better choice. When you cannot be sure the training data covers every category, the docs recommend handle_unknown='infrequent_if_exist' rather than setting the categories by hand. With this setting, an unseen value does not raise an error. Its columns for that feature are all zeros, or the value goes into an infrequent category if infrequent grouping is turned on.

    handle_unknown='infrequent_if_exist': 'from Asia' and 'uses Chrome' were never seen, so their feature columns are all zerospython
    >>> enc = preprocessing.OneHotEncoder(handle_unknown='infrequent_if_exist')
    >>> X = [['male', 'from US', 'uses Safari'], ['female', 'from Europe', 'uses Firefox']]
    >>> enc.fit(X)
    OneHotEncoder(handle_unknown='infrequent_if_exist')
    >>> enc.transform([['female', 'from Asia', 'uses Chrome']]).toarray()
    array([[1., 0., 0., 0., 0., 0.]])

    Checkpoint 3 of 4· Check yourself

    An encoder fitted with handle_unknown='infrequent_if_exist' (no infrequent grouping enabled) is asked to transform a browser value it never saw. What happens?

    Checkpoint 4 of 4· Exam question

    A model is trained with an `OneHotEncoder` fit on a `device_type` column that contained `phone`, `tablet`, and `desktop` during training. In production, an inference batch arrives containing a new value, `wearable`, that the encoder never saw during fitting. The team wants the pipeline to score these rows without crashing. What should they configure on the fitted encoder?

    Sources32

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Integer-coding a nominal feature with OrdinalEncoder is a safe replacement for one-hot encoding with any estimator.Why is that wrong?

      Estimators that expect continuous input treat integer codes as ordered. For nominal labels such as browsers, that order is arbitrary and misleading.

      Covered in Why integer codes are not enough

    2. 2.The only way to cope with categories missing from the training data is to list them all with the categories parameter.Why is that wrong?

      The docs say handle_unknown='infrequent_if_exist' is often the better choice. It encodes unseen values as all zeros, or as an infrequent category, without raising an error.

      Covered in Categories the encoder never saw in training

    3. 3.With SparkR's SparkML GLM you must one-hot encode string features yourself before fitting.Why is that wrong?

      SparkR does this automatically for SparkML GLM, so you do not need to encode those features manually.

      Covered in Fitting OneHotEncoder and reading its output

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually”
      ↩︎ Why integer codes are not enough
      “Under the hood, SparkR uses MLlib to train the model.”
      ↩︎ Fitting OneHotEncoder and reading its output
      “SparkR automatically performs one-hot encoding of categorical features so that it does not need to be done manually”
      ↩︎ Exam trap 3
    2. 2.
      “the set of browsers was ordered arbitrarily”
      ↩︎ Why integer codes are not enough
      “one-of-K, also known as one-hot or dummy encoding”
      ↩︎ Why integer codes are not enough
      “By default, the values each feature can take is inferred automatically from the dataset and can be found in the categories_ attribute”
      ↩︎ Fitting OneHotEncoder and reading its output
      “it can often be better to specify handle_unknown='infrequent_if_exist' instead of setting the categories manually as above”
      ↩︎ Categories the encoder never saw in training
      “transforms each categorical feature with n_categories possible values into n_categories binary features, with one of them 1, and all others 0”
      ↩︎ Key concept
      “would interpret the categories as being ordered, which is often not desired”
      ↩︎ Exam trap 1
      “it can often be better to specify handle_unknown='infrequent_if_exist' instead of setting the categories manually as above”
      ↩︎ Exam trap 2
      “would interpret the categories as being ordered, which is often not desired”
      ↩︎ Prediction
      “no error will be raised but the resulting one-hot encoded columns for this feature will be all zeros”
      ↩︎ Checkpoint
    3. 3.
      “Automatic feature generation processing, like one-hot encoding for categorical features, also occurs during this stage.”
      ↩︎ Fitting OneHotEncoder and reading its output
      “Automatic feature generation processing, like one-hot encoding for categorical features, also occurs during this stage.”
      ↩︎ Categories the encoder never saw in training

    Continue to page 2 of 2

    Tuning OneHotEncoder: drop, Missing Values, Infrequent Categories and Pipelines

    Spotted a mistake, or was something unclear? Tell us.