CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 2 · Lesson 27/48

    Log Transformation: Recognizing Skewed and Multiplicative Data

    Identify scenarios where log scale transformation is appropriate

    9 min read
    2.08% of exam
    5 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain why a logarithm turns equal ratios into equal steps, using the Spark log and log10 functions
    • Identify features that span orders of magnitude or are skewed as candidates for a log scale
    • Recognize residual spread that grows with the prediction as a sign that a variance-stabilizing transform may help
    • Recognize an exponentially generated target as a case where a natural log undoes the curvature

    Key concept

    Log scale transformation — You replace each value with its logarithm, so values that grow by multiplication end up evenly spaced. Doublings or factors of ten each become one equal step, and a long upper tail gets pulled in toward the rest of the data.

    1.What a logarithm does to a scale

    A logarithm asks what power of the base produces a value. That makes it a converter from ratios to steps. The Databricks PySpark log function takes an optional base as its first argument. When you call it with only one argument, it falls back to the natural logarithm. The documentation's own example uses base 2 on the values 1, 2 and 4.

    Spark's log function with an explicit base of 2. The values 1, 2 and 4 come out as 0.0, 1.0 and 2.0.python
    from pyspark.sql import functions as dbf
    df = spark.sql("SELECT * FROM VALUES (1), (2), (4) AS t(value)")
    df.select("*", dbf.log(2.0, df.value)).show()

    That is the first scenario where a log scale fits: a feature whose changes are multiplicative. Think of growth by a factor, a doubling, or a percentage change. On the raw scale, a move from 1 to 2 looks tiny next to a move from 1,000 to 2,000, even though both are the same doubling. On the log scale they are the same size. If the relationship you want a model to learn works in ratios, the log puts it on a scale where equal effects look equal.

    Checkpoint 1 of 5· Check yourself

    A colleague writes dbf.log(df.revenue) with a single argument and assumes the result is base 10. What does Spark actually compute?

    Sources1

    2.Features that span orders of magnitude or are skewed

    The same compression explains the second scenario: a feature whose values run across several orders of magnitude. In Spark's log10 example, 1, 10 and 100 become 0, 1 and 2. A hundredfold range shrinks to a span of two units, and each factor of ten takes the same amount of room.

    Spark's log10 maps 1, 10 and 100 to 0.0, 1.0 and 2.0. Each order of magnitude becomes one unit.python
    from pyspark.sql import functions as dbf
    df = spark.createDataFrame([(1,), (10,), (100,)], ["value"])
    df.select("*", dbf.log10(df.value)).show()

    Picture the histogram of a feature like that. Most rows pile up at the low end and a thin tail stretches far to the right. The log squeezes the big values together much more than the small ones, so the tail gets pulled in. The scikit-learn preprocessing guide describes why reshaping a distribution is worth doing: in many modeling scenarios, normality of the features is desirable. Its family of power transforms aims to map data toward a Gaussian 'in order to stabilize variance and minimize skewness.'

    The guide's worked example is exactly that: Box-Cox turns lognormal samples into normally distributed ones. A skewed feature whose bulk sits low and whose tail runs high is the typical candidate for a log-style transform. Reshaping costs you nothing in ordering. These transforms are monotonic, so the smallest value stays the smallest and the largest stays the largest. Only the distances between values change.

    Checkpoint 2 of 5· Check yourself

    You log-transform a skewed income feature. Which property of the feature is guaranteed to survive the transformation?

    Checkpoint 3 of 5· Exam question

    A data scientist is preparing a `price` feature for a linear regression model predicting home values. The feature is heavily right-skewed, with most homes priced under $500,000 but a small number of luxury properties priced above $5,000,000, stretching the tail. Which transformation is the most appropriate way to prepare this feature before training the model?

    Sources23

    3.Spread that grows with the level, and exponential targets

    The third scenario shows up in diagnostics rather than in a histogram. scikit-learn's guidance on residual plots says that a linear least squares model such as LinearRegression or Ridge expects its residuals to be uncorrelated, centred on zero, and of constant variance. When the true distribution of y given X is Poisson or Gamma, the guide notes that the residual variance is expected to grow with the predicted value. You then see a fan or banana shape instead of an even band. The guide reads such a shape as a hint that non-linear feature engineering or a non-linear model might be useful. A transform that stabilizes variance is one such piece of feature engineering.

    Checkpoint 4 of 5· Check yourself

    A residuals-vs-predicted plot from a LinearRegression model shows the spread of residuals widening steadily as predictions increase. What does this tell you?

    The clearest version of this is a target that was produced by exponentiation. scikit-learn's TransformedTargetRegressor example builds its regression target by passing a linear signal through np.exp. A straight line fits that target poorly. Taking the natural logarithm, which Spark exposes as ln, undoes the exponential and leaves the linear signal underneath. In the example, a LinearRegression on the transformed target scores an R² of 0.67, compared with 0.64 on the raw target.

    How the scikit-learn example creates an exponentially distributed regression targetpython
    >>> y = np.exp( 1 + (y - y.min()) * (4 / (y.max() - y.min())))

    Checkpoint 5 of 5· Exam question

    An ML team is modeling the relationship between weekly marketing spend and website signups. Exploratory plots show that signups grow multiplicatively as spend increases: each additional dollar of spend produces a percentage increase in signups rather than a fixed increase. Which technique best addresses this relationship before fitting a linear regression model?

    Sources45

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Calling Spark's log function with one argument gives a base-10 logarithm.Why is that wrong?

      With a single argument, log computes the natural logarithm. For base 10, use log10 or pass 10 as the base argument.

      Covered in What a logarithm does to a scale

    2. 2.Residuals that fan out as predictions grow are just noise and need no action.Why is that wrong?

      Linear least squares assumes constant residual variance. Spread that grows with the prediction violates that assumption and points toward transforming the target or using a non-linear model.

      Covered in Spread that grows with the level, and exponential targets

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “If there is only one argument, then this takes the natural logarithm of the argument.”
      ↩︎ What a logarithm does to a scale
      “If there is only one argument, then this takes the natural logarithm of the argument.”
      ↩︎ Exam trap 1
    2. 2.
      “Computes the logarithm of the given value in Base 10.”
      ↩︎ Features that span orders of magnitude or are skewed
      “Computes the logarithm of the given value in Base 10.”
      ↩︎ Key concept
    3. 3.
      “aim to map data from any distribution to as close to a Gaussian distribution as possible in order to stabilize variance and minimize skewness”
      ↩︎ Features that span orders of magnitude or are skewed
      “Here is an example of using Box-Cox to map samples drawn from a lognormal distribution to a normal distribution”
      ↩︎ Prediction
      “Both quantile and power transforms are based on monotonic transformations of the features and thus preserve the rank of the values along each feature.”
      ↩︎ Checkpoint
    4. 5.
      “their variance should be constant (homoschedasticity)”
      ↩︎ Spread that grows with the level, and exponential targets
      “it is expected that the variance of the residuals of the optimal model would grow with the predicted value of E[y|X]”
      ↩︎ Exam trap 2

    Continue to page 2 of 2

    Applying Log Transforms Safely: Zeros, Inverses and Alternatives

    Spotted a mistake, or was something unclear? Tell us.