What you will be able to do
- Explain why a logarithm turns equal ratios into equal steps, using the Spark log and log10 functions
- Identify features that span orders of magnitude or are skewed as candidates for a log scale
- Recognize residual spread that grows with the prediction as a sign that a variance-stabilizing transform may help
- Recognize an exponentially generated target as a case where a natural log undoes the curvature
Key concept
Log scale transformation — You replace each value with its logarithm, so values that grow by multiplication end up evenly spaced. Doublings or factors of ten each become one equal step, and a long upper tail gets pulled in toward the rest of the data.
1.What a logarithm does to a scale
A logarithm asks what power of the base produces a value. That makes it a converter from ratios to steps. The Databricks PySpark log function takes an optional base as its first argument. When you call it with only one argument, it falls back to the natural logarithm. The documentation's own example uses base 2 on the values 1, 2 and 4.
from pyspark.sql import functions as dbf
df = spark.sql("SELECT * FROM VALUES (1), (2), (4) AS t(value)")
df.select("*", dbf.log(2.0, df.value)).show()The output is 0.0, 1.0 and 2.0. Each doubling became a step of exactly 1. Two values that are twice as far apart on the raw scale are now the same distance apart, because the transform measures how many times you multiplied, not how much you added.
That is the first scenario where a log scale fits: a feature whose changes are multiplicative. Think of growth by a factor, a doubling, or a percentage change. On the raw scale, a move from 1 to 2 looks tiny next to a move from 1,000 to 2,000, even though both are the same doubling. On the log scale they are the same size. If the relationship you want a model to learn works in ratios, the log puts it on a scale where equal effects look equal.
Checkpoint 1 of 5· Check yourself
A colleague writes dbf.log(df.revenue) with a single argument and assumes the result is base 10. What does Spark actually compute?
With one argument, Spark's log takes the natural logarithm. For base 10 you need log10, or an explicit base as the first argument.
“If there is only one argument, then this takes the natural logarithm of the argument.”Source: docs.databricks.com
Sources1
2.Features that span orders of magnitude or are skewed
The same compression explains the second scenario: a feature whose values run across several orders of magnitude. In Spark's log10 example, 1, 10 and 100 become 0, 1 and 2. A hundredfold range shrinks to a span of two units, and each factor of ten takes the same amount of room.
from pyspark.sql import functions as dbf
df = spark.createDataFrame([(1,), (10,), (100,)], ["value"])
df.select("*", dbf.log10(df.value)).show()Picture the histogram of a feature like that. Most rows pile up at the low end and a thin tail stretches far to the right. The log squeezes the big values together much more than the small ones, so the tail gets pulled in. The scikit-learn preprocessing guide describes why reshaping a distribution is worth doing: in many modeling scenarios, normality of the features is desirable. Its family of power transforms aims to map data toward a Gaussian 'in order to stabilize variance and minimize skewness.'
The guide's worked example is exactly that: Box-Cox turns lognormal samples into normally distributed ones. A skewed feature whose bulk sits low and whose tail runs high is the typical candidate for a log-style transform. Reshaping costs you nothing in ordering. These transforms are monotonic, so the smallest value stays the smallest and the largest stays the largest. Only the distances between values change.
Checkpoint 2 of 5· Check yourself
You log-transform a skewed income feature. Which property of the feature is guaranteed to survive the transformation?
The log, like the power and quantile transforms, is monotonic, so it preserves rank. Distances, means and linear correlations all change, because changing those is the whole point.
“Both quantile and power transforms are based on monotonic transformations of the features and thus preserve the rank of the values along each feature.”Source: scikit-learn.org
Checkpoint 3 of 5· Exam question
A data scientist is preparing a `price` feature for a linear regression model predicting home values. The feature is heavily right-skewed, with most homes priced under $500,000 but a small number of luxury properties priced above $5,000,000, stretching the tail. Which transformation is the most appropriate way to prepare this feature before training the model?
Correct answer: A — Apply a log transformation to the price feature, which compresses the extreme high values and pulls the distribution closer to a normal shape that linear regression assumes.
- A. A log transformation compresses large values proportionally more than small ones, which pulls in the long right tail created by the luxury properties and moves the feature toward the roughly normal shape linear regression works best with.
- B. Min-max scaling only rescales the numeric range of a feature; it applies the same linear stretch to every value, so the right-skewed shape of the price distribution, including the stretched-out tail, remains exactly as it was before scaling.
- C. Converting a continuous price feature into categorical buckets throws away the ordered numeric relationship between prices, which a linear model relies on to estimate a continuous coefficient for the feature.
- D. Deleting the luxury properties removes legitimate high-value homes from the training data rather than reshaping the distribution, which biases the model against ever predicting high prices correctly.
3.Spread that grows with the level, and exponential targets
The third scenario shows up in diagnostics rather than in a histogram. scikit-learn's guidance on residual plots says that a linear least squares model such as LinearRegression or Ridge expects its residuals to be uncorrelated, centred on zero, and of constant variance. When the true distribution of y given X is Poisson or Gamma, the guide notes that the residual variance is expected to grow with the predicted value. You then see a fan or banana shape instead of an even band. The guide reads such a shape as a hint that non-linear feature engineering or a non-linear model might be useful. A transform that stabilizes variance is one such piece of feature engineering.
Checkpoint 4 of 5· Check yourself
A residuals-vs-predicted plot from a LinearRegression model shows the spread of residuals widening steadily as predictions increase. What does this tell you?
Least squares regression assumes residual variance is constant. A spread that grows with the prediction breaks that assumption, which is a typical cue for transforming the target or switching to a non-linear model.
“their variance should be constant (homoschedasticity)”Source: scikit-learn.org
The clearest version of this is a target that was produced by exponentiation. scikit-learn's TransformedTargetRegressor example builds its regression target by passing a linear signal through np.exp. A straight line fits that target poorly. Taking the natural logarithm, which Spark exposes as ln, undoes the exponential and leaves the linear signal underneath. In the example, a LinearRegression on the transformed target scores an R² of 0.67, compared with 0.64 on the raw target.
>>> y = np.exp( 1 + (y - y.min()) * (4 / (y.max() - y.min())))Checkpoint 5 of 5· Exam question
An ML team is modeling the relationship between weekly marketing spend and website signups. Exploratory plots show that signups grow multiplicatively as spend increases: each additional dollar of spend produces a percentage increase in signups rather than a fixed increase. Which technique best addresses this relationship before fitting a linear regression model?
Correct answer: A — Log-transform both the spend and signup features, which converts the multiplicative growth pattern into an approximately linear relationship a model can fit.
- A. When a percentage change in one variable corresponds to a percentage change in another, taking logs of both variables turns that multiplicative relationship into an additive, approximately linear one that a linear model can capture directly.
- B. Standardization centers and rescales a feature using its mean and standard deviation, but it is a linear operation, so it cannot turn a multiplicative, percentage-based growth pattern into a linear one.
- C. Converting spend into quartile bins throws away the continuous, fine-grained growth pattern the model needs to learn and instead forces the relationship into a small number of discrete steps.
- D. Median imputation only fills in missing spend values; it has no effect on the shape of the relationship between spend and signups for the values that are already present.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Calling Spark's log function with one argument gives a base-10 logarithm.Why is that wrong?
With a single argument, log computes the natural logarithm. For base 10, use log10 or pass 10 as the base argument.
Covered in What a logarithm does to a scale
2.Residuals that fan out as predictions grow are just noise and need no action.Why is that wrong?
Linear least squares assumes constant residual variance. Spread that grows with the prediction violates that assumption and points toward transforming the target or using a non-linear model.
Covered in Spread that grows with the level, and exponential targets
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“If there is only one argument, then this takes the natural logarithm of the argument.”
↩︎ What a logarithm does to a scale“If there is only one argument, then this takes the natural logarithm of the argument.”
↩︎ Exam trap 1 - 2.
“Computes the logarithm of the given value in Base 10.”
↩︎ Features that span orders of magnitude or are skewed“Computes the logarithm of the given value in Base 10.”
↩︎ Key concept - 3.https://scikit-learn.org/stable/modules/preprocessing.htmlSecondary source
“aim to map data from any distribution to as close to a Gaussian distribution as possible in order to stabilize variance and minimize skewness”
↩︎ Features that span orders of magnitude or are skewed“Here is an example of using Box-Cox to map samples drawn from a lognormal distribution to a normal distribution”
↩︎ Prediction“Both quantile and power transforms are based on monotonic transformations of the features and thus preserve the rank of the values along each feature.”
↩︎ Checkpoint - 4.
“Returns the natural logarithm of the argument.”
↩︎ Spread that grows with the level, and exponential targets - 5.https://scikit-learn.org/stable/modules/model_evaluation.htmlSecondary source
“their variance should be constant (homoschedasticity)”
↩︎ Spread that grows with the level, and exponential targets“it is expected that the variance of the residuals of the optimal model would grow with the predicted value of E[y|X]”
↩︎ Exam trap 2