CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 2 · Lesson 27/48

    Applying Log Transforms Safely: Zeros, Inverses and Alternatives

    Identify scenarios where log scale transformation is appropriate

    8 min read
    2.08% of exam
    6 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Predict what Spark's ln, log10 and log1p return for zero and negative inputs, and choose log1p for data containing zeros
    • Map predictions from a log-transformed target back to original units with an inverse function
    • Build a log transformation into a scikit-learn pipeline with FunctionTransformer
    • Recognize when a power or quantile transform is a better fit than a plain log

    1.Zeros and negatives: where a plain log breaks

    The logarithm is only defined for positive numbers, and Spark handles bad inputs without any warning. Databricks SQL's log10 returns NULL when its input is less than or equal to 0. The PySpark examples show the same thing: ln over spark.range(10) gives NULL for id 0, and log with base 3 gives NULL for both 0 and -1. A NaN input to log10 comes out as NaN. Nothing fails. You just end up with a column that has picked up new missing values wherever the raw data had a zero.

    Natural log over ids 0 to 9. In the documented output, the row for id 0 is NULL.python
    from pyspark.sql import functions as dbf
    spark.range(10).select("*", dbf.ln('id')).show()

    That matters because many of the features that suit a log are exactly the ones that contain zeros: counts, amounts and durations, where a large share of rows is zero and a few are huge. The usual fix is log1p, which the Databricks SQL reference defines as returning log(1 + expr). Adding one before taking the log moves a zero to log(1), which is 0. The function only returns NULL once its input reaches -1 or below, so any non-negative feature is safe. The same positivity question comes up with scikit-learn's Box-Cox transform: Box-Cox can only be applied to strictly positive data.

    Checkpoint 1 of 5· Check yourself

    A purchase-count feature is heavily skewed, and 40% of its rows are 0. You want a log scale in Spark without creating NULLs. Which function should you use?

    Sources123

    2.Undoing the transform: predictions in original units

    Transforming a feature is a one-way job. The model just sees the new column. Transforming the target is different, because whatever the model predicts comes out on the log scale. A predicted log-price of 2 is not a price. scikit-learn's TransformedTargetRegressor exists for this situation. It transforms the targets y before fitting a regression model, and then maps the predictions back to the original space through an inverse transform. For a simple log, you can give it a pair of functions instead of a transformer object.

    A log transform and its inverse, passed to TransformedTargetRegressor as func and inverse_funcpython
    >>> def func(x):
    ...     return np.log(x)
    >>> def inverse_func(x):
    ...     return np.exp(x)

    The pair has to be a real inverse. By default the regressor checks this on every fit. The documentation shows what goes wrong if you switch the check off with check_inverse=False and supply an identity function as the 'inverse': the R² score drops from 0.67 to -3.02. If you used log1p on the way in, the matching way out is expm1, which Spark documents as computing the exponential of the given value minus one. For features, scikit-learn wraps the same idea in FunctionTransformer, which turns an arbitrary function into a pipeline step.

    Checkpoint 2 of 5· Fill the gap

    Which function completes this scikit-learn sample so that the zero in X maps to 0.0 instead of failing?

    >>> import numpy as np
    >>> from sklearn.preprocessing import FunctionTransformer
    >>> transformer = FunctionTransformer( ? , validate=True)
    >>> X = np.array([[0, 1], [2, 3]])

    Checkpoint 3 of 5· Match them up

    Match each function to its role in a log-transform workflow

    Tap a term, then the definition that fits it.

    Sources45

    3.When a plain log is not enough

    A fixed log is one choice among several monotonic transforms, and scikit-learn offers two more general ones. PowerTransformer provides Box-Cox and Yeo-Johnson, each with a parameter λ fitted by maximum likelihood. By default it also standardizes the output to zero mean and unit variance. QuantileTransformer works from ranks instead. It can push any feature to a uniform or normal shape and is less influenced by outliers than scaling methods. The cost is that it distorts correlations and distances within and across features.

    Choosing a distribution-reshaping transform by input constraint and behaviour
    TransformInput it acceptsWhat it does
    ln / log10 (Spark)Positive values; 0 and below return NULLFixed log; equal ratios become equal steps
    log1p (Spark) / np.log1pValues above -1; 0 maps to 0Natural log of the value plus one
    PowerTransformer, method='box-cox'Strictly positive data onlyFits λ by maximum likelihood to approach a Gaussian
    QuantileTransformerAny values; rank-basedMaps to uniform or normal; less influenced by outliers, distorts correlations

    None of these works on every dataset. The scikit-learn guide notes that on certain distributions the power transforms produce very Gaussian-like results, while on others they are ineffective. Its advice is to visualize the data before and after the transformation. Use the scenarios as a short list of candidates, and let the plot decide.

    Checkpoint 4 of 5· Check yourself

    A feature has extreme outliers, and you want a transform that is barely affected by them. You don't care whether distances between values are preserved. Which option fits best?

    Checkpoint 5 of 5· Exam question

    A dataset of customer support ticket resolution times includes many tickets resolved in under an hour, a long tail of tickets that took several days, and a subset of tickets with a resolution time of exactly zero minutes because they were auto-closed. The team wants to log-transform resolution time to reduce skew before training a model. Which approach correctly handles this feature?

    Sources63

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.You can apply ln or log10 straight to a count feature that contains zeros.Why is that wrong?

      Spark returns NULL for inputs at or below 0, which quietly adds missing values. log1p maps 0 to 0 and is the usual choice for non-negative data with zeros.

      Covered in Zeros and negatives: where a plain log breaks

    2. 2.A regression model trained on a log-transformed target predicts in the target's original units.Why is that wrong?

      Its raw predictions are on the log scale and must go back through the inverse transform, such as exp or expm1. TransformedTargetRegressor does this for you.

      Covered in Undoing the transform: predictions in original units

    3. 3.Any power or log transform will make a skewed feature look Gaussian.Why is that wrong?

      Power transforms work well on some distributions and fail on others, so visualize before and after.

      Covered in When a plain log is not enough

    Practise it for real

    See for yourself in a Databricks notebook why log1p is the safe choice for data that contains zeros, and how expm1 reverses it.

    1. 1.Run spark.range(10).select("*", dbf.ln('id')).show() after from pyspark.sql import functions as dbf.

      Why: This shows how a plain natural log handles a zero.

      You should see: The ln value for id 0 is NULL, and id 1 gives 0.0.

    2. 2.Run the same query with dbf.log1p('id') in place of dbf.ln('id').

      Why: log1p computes the log of the value plus one, so 0 becomes log(1).

      You should see: No NULLs. Id 0 gives 0.0, and id 1 gives the same value that ln gave for id 2.

    3. 3.Wrap the log1p column in dbf.expm1(...) and show the result next to id.

      Why: expm1 is the exponential minus one, the inverse you need to map predictions back to original units.

      You should see: The round-tripped column matches id, up to floating-point precision.

    Stuck? Get a nudge

    If the round-trip doesn't match, check that you used expm1 and not exp. exp alone leaves every value one too high.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “If expr is less than or equal to -1 the result is NULL.”
      ↩︎ Checkpoint
    2. 2.
      “If expr is less than or equal to 0 the result is NULL.”
      ↩︎ Zeros and negatives: where a plain log breaks
      “If expr is less than or equal to 0 the result is NULL.”
      ↩︎ Exam trap 1
    3. 3.
      “Box-Cox can only be applied to strictly positive data.”
      ↩︎ Zeros and negatives: where a plain log breaks
      “This highlights the importance of visualizing the data before and after transformation.”
      ↩︎ When a plain log is not enough
      “It does, however, distort correlations and distances within and across features.”
      ↩︎ When a plain log is not enough
      “the power transforms achieve very Gaussian-like results, but with others, they are ineffective.”
      ↩︎ Exam trap 3
      “a quantile transform smooths out unusual distributions and is less influenced by outliers than scaling methods”
      ↩︎ Checkpoint
    4. 5.
      “TransformedTargetRegressor transforms the targets y before fitting a regression model.”
      ↩︎ Undoing the transform: predictions in original units
      “The predictions are mapped back to the original space via an inverse transform.”
      ↩︎ Exam trap 2
      “The predictions are mapped back to the original space via an inverse transform.”
      ↩︎ Checkpoint
    5. 6.
      “Computes the natural logarithm of the given value plus one.”
      ↩︎ When a plain log is not enough

    Spotted a mistake, or was something unclear? Tell us.