What you will be able to do
- Predict what Spark's ln, log10 and log1p return for zero and negative inputs, and choose log1p for data containing zeros
- Map predictions from a log-transformed target back to original units with an inverse function
- Build a log transformation into a scikit-learn pipeline with FunctionTransformer
- Recognize when a power or quantile transform is a better fit than a plain log
1.Zeros and negatives: where a plain log breaks
The logarithm is only defined for positive numbers, and Spark handles bad inputs without any warning. Databricks SQL's log10 returns NULL when its input is less than or equal to 0. The PySpark examples show the same thing: ln over spark.range(10) gives NULL for id 0, and log with base 3 gives NULL for both 0 and -1. A NaN input to log10 comes out as NaN. Nothing fails. You just end up with a column that has picked up new missing values wherever the raw data had a zero.
from pyspark.sql import functions as dbf
spark.range(10).select("*", dbf.ln('id')).show()That matters because many of the features that suit a log are exactly the ones that contain zeros: counts, amounts and durations, where a large share of rows is zero and a few are huge. The usual fix is log1p, which the Databricks SQL reference defines as returning log(1 + expr). Adding one before taking the log moves a zero to log(1), which is 0. The function only returns NULL once its input reaches -1 or below, so any non-negative feature is safe. The same positivity question comes up with scikit-learn's Box-Cox transform: Box-Cox can only be applied to strictly positive data.
Checkpoint 1 of 5· Check yourself
A purchase-count feature is heavily skewed, and 40% of its rows are 0. You want a log scale in Spark without creating NULLs. Which function should you use?
log1p computes log(1 + x), so a 0 maps to 0, and it only returns NULL for inputs of -1 or below. ln, log10 and log with any base all return NULL at 0.
“If expr is less than or equal to -1 the result is NULL.”Source: docs.databricks.com
2.Undoing the transform: predictions in original units
Transforming a feature is a one-way job. The model just sees the new column. Transforming the target is different, because whatever the model predicts comes out on the log scale. A predicted log-price of 2 is not a price. scikit-learn's TransformedTargetRegressor exists for this situation. It transforms the targets y before fitting a regression model, and then maps the predictions back to the original space through an inverse transform. For a simple log, you can give it a pair of functions instead of a transformer object.
>>> def func(x):
... return np.log(x)
>>> def inverse_func(x):
... return np.exp(x)The pair has to be a real inverse. By default the regressor checks this on every fit. The documentation shows what goes wrong if you switch the check off with check_inverse=False and supply an identity function as the 'inverse': the R² score drops from 0.67 to -3.02. If you used log1p on the way in, the matching way out is expm1, which Spark documents as computing the exponential of the given value minus one. For features, scikit-learn wraps the same idea in FunctionTransformer, which turns an arbitrary function into a pipeline step.
Checkpoint 2 of 5· Fill the gap
Which function completes this scikit-learn sample so that the zero in X maps to 0.0 instead of failing?
>>> import numpy as np
>>> from sklearn.preprocessing import FunctionTransformer
>>> transformer = FunctionTransformer( ? , validate=True)
>>> X = np.array([[0, 1], [2, 3]])The documented output turns 0 into 0.0 and 1 into 0.69314718, which is ln(2). That only happens with log1p, which adds one before taking the natural log.
Source: scikit-learn.orgCheckpoint 3 of 5· Match them up
Match each function to its role in a log-transform workflow
Tap a term, then the definition that fits it.
log1p and expm1 are each other's inverses. TransformedTargetRegressor applies a forward and inverse pair around a regressor, and FunctionTransformer turns a function like np.log1p into a transformer.
“The predictions are mapped back to the original space via an inverse transform.”Source: scikit-learn.org
3.When a plain log is not enough
A fixed log is one choice among several monotonic transforms, and scikit-learn offers two more general ones. PowerTransformer provides Box-Cox and Yeo-Johnson, each with a parameter λ fitted by maximum likelihood. By default it also standardizes the output to zero mean and unit variance. QuantileTransformer works from ranks instead. It can push any feature to a uniform or normal shape and is less influenced by outliers than scaling methods. The cost is that it distorts correlations and distances within and across features.
| Transform | Input it accepts | What it does |
|---|---|---|
| ln / log10 (Spark) | Positive values; 0 and below return NULL | Fixed log; equal ratios become equal steps |
| log1p (Spark) / np.log1p | Values above -1; 0 maps to 0 | Natural log of the value plus one |
| PowerTransformer, method='box-cox' | Strictly positive data only | Fits λ by maximum likelihood to approach a Gaussian |
| QuantileTransformer | Any values; rank-based | Maps to uniform or normal; less influenced by outliers, distorts correlations |
None of these works on every dataset. The scikit-learn guide notes that on certain distributions the power transforms produce very Gaussian-like results, while on others they are ineffective. Its advice is to visualize the data before and after the transformation. Use the scenarios as a short list of candidates, and let the plot decide.
Checkpoint 4 of 5· Check yourself
A feature has extreme outliers, and you want a transform that is barely affected by them. You don't care whether distances between values are preserved. Which option fits best?
Because the quantile transform is rank-based, it is less influenced by outliers than scaling methods. That robustness costs you the original correlations and distances.
“a quantile transform smooths out unusual distributions and is less influenced by outliers than scaling methods”Source: scikit-learn.org
Checkpoint 5 of 5· Exam question
A dataset of customer support ticket resolution times includes many tickets resolved in under an hour, a long tail of tickets that took several days, and a subset of tickets with a resolution time of exactly zero minutes because they were auto-closed. The team wants to log-transform resolution time to reduce skew before training a model. Which approach correctly handles this feature?
Correct answer: A — Apply `log1p` (log of one plus the value) to resolution time, which is defined at zero and still compresses the long right tail of multi-day tickets.
- A. The `log1p` function adds one before taking the log, which makes it well-defined at zero while still compressing the large values in the long tail, so the auto-closed tickets do not need to be dropped or altered.
- B. The natural logarithm is undefined at zero, so applying it directly would produce an error or an invalid result for every auto-closed ticket rather than leaving those rows unchanged.
- C. A square root transformation is defined at zero, but it compresses large values far less aggressively than a logarithm, so it does not reduce a heavy right tail as strongly as the log-based approach.
- D. Dropping the zero-minute tickets removes a real and fairly common outcome from the training data just to make a direct logarithm computable, losing information the model could otherwise learn from.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.You can apply ln or log10 straight to a count feature that contains zeros.Why is that wrong?
Spark returns NULL for inputs at or below 0, which quietly adds missing values. log1p maps 0 to 0 and is the usual choice for non-negative data with zeros.
2.A regression model trained on a log-transformed target predicts in the target's original units.Why is that wrong?
Its raw predictions are on the log scale and must go back through the inverse transform, such as exp or expm1. TransformedTargetRegressor does this for you.
Covered in Undoing the transform: predictions in original units
3.Any power or log transform will make a skewed feature look Gaussian.Why is that wrong?
Power transforms work well on some distributions and fail on others, so visualize before and after.
Covered in When a plain log is not enough
Practise it for real
See for yourself in a Databricks notebook why log1p is the safe choice for data that contains zeros, and how expm1 reverses it.
1.Run
spark.range(10).select("*", dbf.ln('id')).show()afterfrom pyspark.sql import functions as dbf.Why: This shows how a plain natural log handles a zero.
You should see: The ln value for id 0 is NULL, and id 1 gives 0.0.
2.Run the same query with
dbf.log1p('id')in place ofdbf.ln('id').Why: log1p computes the log of the value plus one, so 0 becomes log(1).
You should see: No NULLs. Id 0 gives 0.0, and id 1 gives the same value that ln gave for id 2.
3.Wrap the log1p column in
dbf.expm1(...)and show the result next toid.Why: expm1 is the exponential minus one, the inverse you need to map predictions back to original units.
You should see: The round-tripped column matches id, up to floating-point precision.
Stuck? Get a nudge
If the round-trip doesn't match, check that you used expm1 and not exp. exp alone leaves every value one too high.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Returns log(1 + expr).”
↩︎ Zeros and negatives: where a plain log breaks“If expr is less than or equal to -1 the result is NULL.”
↩︎ Checkpoint - 2.
“If expr is less than or equal to 0 the result is NULL.”
↩︎ Zeros and negatives: where a plain log breaks“If expr is less than or equal to 0 the result is NULL.”
↩︎ Exam trap 1 - 3.https://scikit-learn.org/stable/modules/preprocessing.htmlSecondary source
“Box-Cox can only be applied to strictly positive data.”
↩︎ Zeros and negatives: where a plain log breaks“This highlights the importance of visualizing the data before and after transformation.”
↩︎ When a plain log is not enough“It does, however, distort correlations and distances within and across features.”
↩︎ When a plain log is not enough“the power transforms achieve very Gaussian-like results, but with others, they are ineffective.”
↩︎ Exam trap 3“a quantile transform smooths out unusual distributions and is less influenced by outliers than scaling methods”
↩︎ Checkpoint - 4.
“Computes the exponential of the given value minus one.”
↩︎ Undoing the transform: predictions in original units - 5.https://scikit-learn.org/stable/modules/compose.htmlSecondary source
“TransformedTargetRegressor transforms the targets y before fitting a regression model.”
↩︎ Undoing the transform: predictions in original units“The predictions are mapped back to the original space via an inverse transform.”
↩︎ Exam trap 2“The predictions are mapped back to the original space via an inverse transform.”
↩︎ Checkpoint - 6.
“Computes the natural logarithm of the given value plus one.”
↩︎ When a plain log is not enough