CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 29/48

    Mitigating Class Imbalance: Resampling, SMOTE and Class Weights

    Identify methods to mitigate data imbalance in training data

    14 min read
    2.08% of exam
    4 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Recognise when a skewed label distribution will distort what a classifier learns
    • Undersample a majority class with PySpark sampleBy and describe random oversampling and SMOTE
    • Explain the two uses of class weights, including how Databricks AutoML pairs downsampling with compensating weights
    • Apply imbalance mitigation to the training data only, and evaluate on the true class distribution

    Key concept

    Rebalance the training data, not the evaluation data — Imbalance mitigation (undersampling, oversampling, SMOTE or class weights) changes what the model sees while it learns. Validation and test data must keep the real class mix so the metrics describe production behaviour.

    1.Why a skewed label column is a training problem

    A dataset is imbalanced when one label value vastly outnumbers another: fraud against legitimate transactions, churners against loyal customers, defective parts against good ones. Most learning algorithms minimise an average loss over all rows. When 95 rows out of 100 belong to one class, the cheapest way to get a low average loss is to lean heavily toward that class. The model can then reach high accuracy while barely learning the minority class, which is usually the one the business cares about.

    There are two levers for fixing this, and the exam expects you to know both. You can change the data the model trains on: remove majority rows (undersampling), add minority rows (oversampling, or synthetic rows with SMOTE). Or you can change how much each row counts in the loss (class weights). Databricks AutoML uses a combination of both levers. When it detects an imbalanced classification dataset, it tries to reduce the imbalance of the training dataset by downsampling the major class(es) and adding class weights. The rest of this lesson takes each lever in turn, then covers the rule that ties them together: rebalance training data only.

    Checkpoint 1 of 8· Check yourself

    When Databricks AutoML detects an imbalanced classification dataset, how does it try to reduce the imbalance of the training data?

    Sources1

    2.Undersampling the majority class with sampleBy

    Undersampling (also called downsampling) throws away a share of the majority-class rows so the classes sit closer in size. It is cheap: the training set gets smaller, so training gets faster. The cost is information. Every discarded row is an example the model never sees, so aggressive undersampling on a small dataset can hurt more than the imbalance did.

    On Spark DataFrames the natural tool is stratified sampling with sampleBy. You name the label column as the stratum column and give a sampling fraction for each label value. To undersample, give the majority label a fraction below 1.0 and the minority label a fraction of 1.0 so every rare row is kept.

    sampleBy draws a different fraction from each stratum; here key 0 keeps about 10% and key 1 about 20%python
    from pyspark.sql import functions as sf
    dataset = spark.range(0, 100, 1, 5).select((sf.col("id") % 3).alias("key"))
    sampled = dataset.sampleBy("key", fractions={0: 0.1, 1: 0.2}, seed=0)
    sampled.groupBy("key").count().orderBy("key").show()

    Look at what happened to key 2 in that example. It was not listed in fractions, and it does not appear in the output at all. That is the documented behaviour: a stratum you leave out of the dictionary gets a fraction of zero. When you undersample, you must list the minority label explicitly with 1.0. If you only write down the class you want to shrink, you delete the class you were trying to protect. Note too that sampleBy samples without replacement, so it can only shrink a stratum, never grow it.

    Checkpoint 2 of 8· Check yourself

    You call df.sampleBy("label", fractions={0: 0.05}, seed=42) on a binary dataset where label 1 is the rare class. What happens to the label-1 rows?

    Sources2

    3.Oversampling and SMOTE: growing the minority class

    Oversampling goes the other way: instead of removing majority rows, you add minority rows. No information is thrown away, but the training set grows and the model can overfit to the repeated rows.

    Random oversampling simply duplicates existing minority rows, which means sampling with replacement. In PySpark, DataFrame.sample has a withReplacement flag for this; its default is False. A common pattern is to filter the minority rows, sample them with replacement, and union the result back onto the training DataFrame. Be aware that sample is approximate: it does not guarantee exactly the requested fraction of rows.

    Checkpoint 3 of 8· Fill the gap

    Which parameter makes this PySpark call sample rows with replacement, the building block of random oversampling?

    df.sample( ? =True, fraction=0.5, seed=3).count()

    SMOTE (Synthetic Minority Over-sampling Technique) is the method the exam names most often. Instead of copying minority rows, it creates new synthetic ones. For a minority example it picks one of its nearest minority-class neighbours and generates a new point somewhere on the line between the two in feature space. Because the new rows are not exact duplicates, SMOTE usually overfits less than random oversampling. It does assume numeric features where interpolation makes sense, so categorical columns need care. In Python, SMOTE comes from the open-source imbalanced-learn package rather than scikit-learn itself.

    Whatever the method, resample after you split, and only inside the training portion. If you oversample or apply SMOTE before splitting, copies or close interpolations of the same minority row land in both training and validation data, and the validation score is inflated by leakage. Under cross-validation this means resampling inside each training fold, not once up front.

    Checkpoint 4 of 8· Exam question

    A data scientist on Databricks is training a `LogisticRegression` model on a fraud dataset where only 2% of transactions are fraudulent. The team wants to correct for this imbalance during training itself, without discarding any rows and without duplicating any existing rows. Which approach best satisfies these constraints?

    Sources3

    4.Class weights: changing how much each row counts

    Class weights leave the rows alone and change the loss instead. Each row's error is multiplied by the weight of its class, so a mistake on a heavily weighted class costs more. In scikit-learn, many classifiers accept a class_weight parameter, and setting it to 'balanced' weights each class inversely to its frequency. That is the most common use of class weights: upweight the rare class so the model can no longer ignore it, with no rows added or removed.

    Databricks AutoML shows a second, subtler use. It downsamples the majority class and then gives that class a compensating weight, so the model trains faster on fewer rows but still learns the true class balance.

    The AutoML worked example: downsampling class A, then weighting it back up
    QuantityClass A (majority)Class B (minority)
    Rows in original training data955
    Rows after downsampling705
    Downsampling ratio70/95 = 0.7361 (not downsampled)
    Class weight used in training1/0.736 = 1.3581

    Why weight the majority class back up? Undersampling alone shifts the class mix the model sees, which shifts its predicted probabilities: a model trained on 70:5 believes class B is more common than it really is. The documentation gives the purpose as making sure the final model is correctly calibrated, with an output probability distribution that matches the input. So the two uses of class weights pull in different directions. Generic minority upweighting deliberately biases the model toward the rare class to catch more of it. AutoML's compensating weights restore the original balance after a speed-motivated downsample. Read exam scenarios carefully to see which goal is in play.

    Checkpoint 5 of 8· Put it in order

    Put AutoML's imbalance handling for a classification training set in order

    1. 1.Downsample the majority class(es) and keep the minority rows
    2. 2.Pass the class weights as a parameter during model training
    3. 3.Detect that the training dataset is imbalanced
    4. 4.Compute class weights inversely related to each class's downsampling ratio

    Checkpoint 6 of 8· Exam question

    A team building a churn model on Databricks decides to apply SMOTE to the minority (churned) class before training. A reviewer asks how SMOTE actually generates its additional training rows. Which description is accurate?

    Sources1

    5.Rebalance training data only; evaluate on the real distribution

    Every technique above changes the data the model learns from. None of them should touch the data you measure it on. Production traffic will arrive with the real, skewed class mix. A test set that has been undersampled or padded with SMOTE rows reports how the model would do in a world that does not exist. Databricks AutoML follows this rule explicitly: it balances only the training split, so model performance is always evaluated on the non-enriched dataset with the true input class distribution.

    The partner habit is to split so that every split still contains the rare class. A plain random split or a plain K-fold can produce a validation fold with no occurrence of a rare class, which leaves metrics such as ROC AUC undefined. Stratification fixes this by keeping the class proportions roughly the same in each split. AutoML's default random split is stratified for classification. When it has to sample a large dataset, it uses sampleBy so that the target label distribution is preserved. In scikit-learn, StratifiedKFold and StratifiedShuffleSplit do the same for cross-validation. Note that stratification preserves the imbalance in each split. It does not fix it. It is what makes honest evaluation possible after you have mitigated the imbalance in training.

    Checkpoint 7 of 8· Check yourself

    A team applies SMOTE to their full labelled dataset, then splits it 80/20 into train and test. Test recall on the minority class looks excellent. What is the main problem?

    Checkpoint 8 of 8· Exam question

    An engineer randomly undersamples the majority class of a 500,000-row imbalanced dataset down to match a 3,000-row minority class before training a classifier. What is the most likely downside of this specific approach?

    Sources14

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.After rebalancing the training set, you should rebalance the validation and test sets the same way for consistency.Why is that wrong?

      Only the training data is rebalanced. Validation and test data keep the true class distribution so metrics reflect real-world performance.

      Covered in Rebalance training data only; evaluate on the real distribution

    2. 2.Class weights always upweight the minority class.Why is that wrong?

      In AutoML's scheme the weights compensate for downsampling: the downsampled majority class A gets weight 1.358 and minority class B stays at 1, so the model stays calibrated.

      Covered in Class weights: changing how much each row counts

    3. 3.With sampleBy you only need to list the majority class you want to shrink; other classes are kept as they are.Why is that wrong?

      Any stratum missing from the fractions dictionary is sampled at zero, so omitting the minority label removes it entirely.

      Covered in Undersampling the majority class with sampleBy

    4. 4.A plain random K-fold split is fine for rare-class problems as long as the dataset is large.Why is that wrong?

      Unstratified splits can produce folds with no rare-class rows, which leaves metrics undefined. Stratified splitters keep class frequencies roughly constant per fold.

      Covered in Rebalance training data only; evaluate on the real distribution

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “it tries to reduce the imbalance of the training dataset by downsampling the major class(es) and adding class weights”
      ↩︎ Why a skewed label column is a training problem
      “AutoML uses class weights that are inversely related to the degree by which a given class is downsampled”
      ↩︎ Class weights: changing how much each row counts
      “To ensure that the final model is correctly calibrated”
      ↩︎ Class weights: changing how much each row counts
      “AutoML only balances the training dataset and does not balance the test and validation datasets.”
      ↩︎ Rebalance training data only; evaluate on the real distribution
      “For classification, a stratified random split ensures that each class is adequately represented in the training, validation, and test sets.”
      ↩︎ Rebalance training data only; evaluate on the real distribution
      “For classification problems, AutoML uses the PySpark sampleBy method for stratified sampling to preserve the target label distribution.”
      ↩︎ Rebalance training data only; evaluate on the real distribution
      “AutoML only balances the training dataset and does not balance the test and validation datasets.”
      ↩︎ Key concept
      “AutoML only balances the training dataset and does not balance the test and validation datasets.”
      ↩︎ Exam trap 1
      “AutoML uses class weights that are inversely related to the degree by which a given class is downsampled”
      ↩︎ Exam trap 2
      “it tries to reduce the imbalance of the training dataset by downsampling the major class(es) and adding class weights”
      ↩︎ Checkpoint
      “Doing so ensures that the model performance is always evaluated on the non-enriched dataset with the true input class distribution.”
      ↩︎ Checkpoint
    2. 2.
      “Returns a stratified sample without replacement based on the fraction given on each stratum.”
      ↩︎ Undersampling the majority class with sampleBy
      “If a stratum is not specified, we treat its fraction as zero.”
      ↩︎ Undersampling the majority class with sampleBy
      “If a stratum is not specified, we treat its fraction as zero.”
      ↩︎ Exam trap 3
    3. 3.
      “Sample with replacement or not (default False).”
      ↩︎ Oversampling and SMOTE: growing the minority class
      “This is not guaranteed to provide exactly the fraction specified of the total count of the given DataFrame.”
      ↩︎ Oversampling and SMOTE: growing the minority class
    4. 4.
      “splitters such as StratifiedKFold and StratifiedShuffleSplit implement stratified sampling to ensure that relative class frequencies are approximately preserved in each fold.”
      ↩︎ Rebalance training data only; evaluate on the real distribution
      “cross-validation splitting can generate train or validation folds without any occurrence of a particular class.”
      ↩︎ Exam trap 4

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.