What you will be able to do
- Recognise that a regression model trained on a log-transformed label returns predictions on the log scale
- Pick the correct inverse function (exp for ln, expm1 for log1p, pow for log10) to bring predictions back to the original units
- Explain why RMSE, MAE and similar metrics must be computed on exponentiated predictions against the original label, not on log values
- Interpret predictions and effects from a log-target model in the original business units
Key concept
Prediction scale follows target scale — A model returns predictions in whatever scale its label had during training. If you trained on log(y), every prediction is a log value, so you must apply the inverse (exponentiate) before you compare it with real labels, compute an error metric, or report it to anyone.
1.A log-target model predicts logs, not prices
Regression targets such as house prices, revenue, salaries or transaction amounts are usually right-skewed: most values are modest and a few are huge. A common fix is to train on the logarithm of the label instead of the raw value. The log pulls in the long tail, makes the residuals more symmetric, and stops a handful of extreme rows from dominating a squared-error loss. In Databricks SQL and PySpark this usually means ln (or log with a single argument, which is the same thing), applied to the label column before training.
The catch is simple, but the exam keeps testing it. The model has no idea the column used to be a price. It learned to predict whatever numbers you gave it, so if the label was ln(price), then predict returns ln(price). For a house that sells for about $200,000, the natural log is about 12.2, and 12.2 is what comes out of the model. Nothing in the training loop converts it back for you. The exception is a wrapper built for this job, such as scikit-learn's TransformedTargetRegressor, which applies the inverse automatically.
Checkpoint 1 of 8· Check yourself
A linear regression was trained on a label column computed as ln(price). For one test row, model.predict returns 12.2. What is the correct reading of this value?
The model predicts on the scale it was trained on, which here is the natural log. Exponentiating (e to the power of the prediction) returns it to dollars.
“Returns e to the power of expr.”Source: docs.databricks.com
2.Undoing the transform: match the inverse to the forward function
To exponentiate correctly, you need to know exactly which log was applied, because each forward function has its own inverse. The natural log ln(x) is undone by exp(x), which is e raised to the prediction. Many pipelines use log1p(x) instead, which is ln(1 + x). It exists because a plain log fails on zero: Databricks returns NULL when the argument is less than or equal to 0, and labels such as sales or counts often contain zeros. The inverse of log1p is expm1, which is e^x − 1. If you apply plain exp to a log1p prediction, every value comes out one unit too high. That error is easy to miss when values are large and very visible when they are near zero.
| Forward transform on the label | Inverse to apply to predictions | When you see it |
|---|---|---|
| ln(col) / log(col) | exp(col) | Strictly positive skewed labels such as prices |
| log1p(col) | expm1(col) | Labels that can be exactly 0, such as counts or sales |
| log10(col) | pow(10, col) | Logarithm in Base 10, used less often in ML pipelines |
| log2(col) | pow(2, col) | Base-2 log; invert by raising 2 to the prediction |
from pyspark.sql import functions as dbf
df = spark.sql("SELECT id AS value FROM RANGE(5)")
df.select("*", dbf.expm1(df.value)).show()Checkpoint 2 of 8· Match them up
Match each forward transform applied to the label with the function that brings predictions back to original units
Tap a term, then the definition that fits it.
Each inverse must undo its own forward function exactly. log1p adds one before taking the log, so its inverse, expm1, subtracts one after exponentiating.
“Computes the exponential of the given value minus one.”Source: docs.databricks.com
Checkpoint 3 of 8· Fill the gap
The label was transformed with ln. Which PySpark function completes this sample so that it returns the exponential of the column?
from pyspark.sql import functions as dbf
df = spark.sql("SELECT id AS value FROM RANGE(5)")
df.select("*", dbf. ? (df.value)).show()exp computes e to the power of the value, which exactly inverts the natural log. expm1 would be correct only if the label had been transformed with log1p.
Source: docs.databricks.comCheckpoint 4 of 8· Exam question
A data scientist trained a regression model to predict home sale prices by first transforming the target using `np.log1p(price)`, since the raw distribution was heavily right-skewed. The model's predictions come out in this transformed scale. Before reporting RMSE to stakeholders in dollar terms, what should the data scientist do?
Correct answer: A — Apply `np.expm1` to both the predicted and actual log-transformed values to reverse the transform, then compute RMSE on the resulting dollar-scale numbers for stakeholders.
- A. This is correct: `np.expm1` is the exact inverse of `np.log1p`, so applying it to both the predictions and the actuals converts them back to dollars before RMSE is computed, giving an error figure stakeholders can interpret in real price terms.
- B. Reporting the error on the log-transformed scale gives a number in log-dollar units that does not correspond to any real price difference, which makes it meaningless to a business audience expecting dollar figures.
- C. Using `np.exp` instead of `np.expm1` produces predictions that are all offset by roughly one dollar relative to the correct inverse, and mixing exponentiated predictions with raw actuals still leaves the units inconsistent with how the model was trained.
- D. Multiplying by the mean of the raw target has no mathematical relationship to undoing a log transform and would produce an arbitrary, incorrect approximation rather than a valid dollar-scale error metric.
3.Compute evaluation metrics after exponentiating, not before
Regression metrics such as RMSE and MAE are built from the difference between a prediction and a true label. The Databricks glm tutorial below shows this directly: the error column is price - prediction, and RMSE is the square root of the mean squared error. That number is in dollars only because both columns are in dollars. If the model had been trained on ln(price), its prediction column would hold log values. Subtracting those from raw prices produces meaningless numbers, and subtracting them from logged prices produces an error in log units, which is not what a business stakeholder or an exam question asking for RMSE in the original units wants.
errors <- select(predictions, predictions$price, predictions$prediction, alias(predictions$price - predictions$prediction, "error"))
display(errors)
# Calculate RMSE
head(select(errors, alias(sqrt(sum(errors$error^2 , na.rm = TRUE) / nrow(errors)), "RMSE")))So the order of operations is fixed. Transform the label, train, predict (getting log values), exponentiate the predictions, and only then compute the metric against the original, untransformed labels. Do not take a shortcut by computing RMSE on the log scale and exponentiating the result. RMSE is a nonlinear summary over all rows, so exp(RMSE_log) is not the dollar RMSE. The exponentiation has to happen row by row, on the predictions, before any averaging. scikit-learn's TransformedTargetRegressor builds this into the estimator: it fits on the transformed target and maps predictions back with the inverse transform, so predict and score already work in the original units.
Checkpoint 5 of 8· Put it in order
Put the steps for evaluating a log-target regression model in the correct order
- 1.Fit the regression model on the transformed target
- 2.Generate predictions, which come out on the log scale
- 3.Map the predictions back to the original space with the inverse transform (exponentiate)
- 4.Apply the log transform to the target y
The transform comes before fitting, and the inverse is applied to the predictions before they are used in any comparison with real labels.
“TransformedTargetRegressor transforms the targets y before fitting a regression model.”Source: scikit-learn.org
Checkpoint 6 of 8· Exam question
A machine learning pipeline predicts subscription lifetime value (LTV) after applying a log transformation to the training target to reduce skew. A product manager asks the data science team to add the model's raw predictions directly to a customer-facing dashboard showing expected dollar value per customer. What should the team do before displaying these numbers?
Correct answer: A — Exponentiate the model's log-scale predictions using the inverse of whatever transform was applied during training, so the dashboard shows dollar values directly.
- A. This is correct: reversing the training-time log transform with its matching inverse function converts the predictions back into dollar units, which is what a dashboard labeled as showing expected dollar value requires.
- B. Preserving rank order does not satisfy the dashboard's requirement, since the request is for dollar-denominated values a product manager can read directly, not just a relative ordering of customers.
- C. Dividing by the standard deviation of the target is a standardization operation, not an inverse of a log transform, and would produce values with no meaningful dollar interpretation.
- D. Rounding a log-scale number does not convert it into dollars; log-transformed values are numerically much smaller than the original target and would badly understate the true expected value.
Checkpoint 7 of 8· Exam question
Two candidate regression models were trained to predict a heavily skewed shipping-cost target after log-transforming it. During evaluation, one engineer reports Model A's RMSE computed directly on the log-transformed predictions, while another reports Model B's RMSE after exponentiating predictions back to the original dollar scale. The team wants to pick the better model using these two RMSE values. What is the problem with this comparison?
Correct answer: A — The two RMSE values are computed in different units, log-dollars versus dollars, and cannot be compared directly; both models must be exponentiated to the same scale first.
- A. This is correct: RMSE computed on log-transformed values and RMSE computed on exponentiated, original-scale values are expressed in different units, so a fair comparison requires putting both models' errors on the same scale first.
- B. RMSE remains a usable metric for skewed targets as long as it is computed consistently in the same units for every model; the issue here is inconsistent units between the two reports, not an inherent flaw in RMSE itself.
- C. Log-scale errors are not guaranteed to be smaller in a way that predicts the original-scale ranking, so assuming Model A is better without exponentiating Model A's predictions is an unjustified conclusion.
- D. Using the same transform during training does not make evaluation metrics comparable when one model's error is reported before exponentiation and the other's is reported after; the units still differ between the two reports.
4.Interpreting log-target predictions and effects
The same rule applies when you explain a model, not just when you score it. A prediction of 12.2 means nothing to a pricing team, while about $199,000 does. Every number that leaves the notebook (a dashboard value, a forecast in a report, a threshold for an alert) should be exponentiated with the inverse that matches the forward transform. In Databricks SQL that is exp, which returns e to the power of its argument, and in PySpark it is dbf.exp (or dbf.expm1 for log1p labels).
Checkpoint 8 of 8· Check yourself
A model was trained on ln(price) and a Databricks SQL dashboard must show its predictions in dollars. Which function should be applied to the prediction column?
exp returns e to the power of its argument, which exactly undoes the natural log. expm1 would only fit a log1p label, and applying another log would move further from dollars.
“Returns e to the power of expr.”Source: docs.databricks.com
Effects change shape as well. With a log-transformed label, a linear model's coefficient no longer means 'adds this many dollars'. It acts multiplicatively on the original scale: a one-unit increase in a feature multiplies the predicted value by exp(coefficient). For small coefficients this is roughly a percentage change of 100 × coefficient. Reading a log-scale coefficient as an additive dollar amount is the interpretation version of the same mistake.
Each extra bathroom multiplies the predicted price by exp(0.05) ≈ 1.051, which is roughly a 5% increase. It does not add $0.05 or 0.05 log-dollars.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.An RMSE computed between log predictions and log labels can be reported as the model's error in dollars.Why is that wrong?
That RMSE is in log units. A dollar RMSE needs predictions mapped back to the original space and compared with the original labels.
Covered in Compute evaluation metrics after exponentiating, not before
2.You can compute RMSE on the log scale and then exponentiate the RMSE to get the original-scale error.Why is that wrong?
The metric must measure the distance between original-scale predictions and original labels, so exponentiate each prediction first and then compute the metric.
Covered in Compute evaluation metrics after exponentiating, not before
3.Plain exp correctly inverts any log transform, including log1p.Why is that wrong?
log1p is ln(1 + x), so its inverse is expm1 (the exponential minus one). Using exp leaves every prediction one unit too high.
Covered in Undoing the transform: match the inverse to the forward function
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Returns the natural logarithm (base e) of expr.”
↩︎ A log-target model predicts logs, not prices - 2.
“log(expr) is a synonym for ln(expr).”
↩︎ A log-target model predicts logs, not prices“If base or expr are less than or equal to 0 the result is NULL.”
↩︎ Undoing the transform: match the inverse to the forward function - 3.https://scikit-learn.org/stable/modules/compose.htmlSecondary source
“TransformedTargetRegressor deals with transforming the target (i.e. log-transform y).”
↩︎ A log-target model predicts logs, not prices“The predictions are mapped back to the original space via an inverse transform.”
↩︎ Compute evaluation metrics after exponentiating, not before“TransformedTargetRegressor deals with transforming the target (i.e. log-transform y).”
↩︎ Key concept“The predictions are mapped back to the original space via an inverse transform.”
↩︎ Exam trap 1“TransformedTargetRegressor transforms the targets y before fitting a regression model.”
↩︎ Checkpoint - 4.
“Returns e to the power of expr.”
↩︎ Undoing the transform: match the inverse to the forward function“Returns e to the power of expr.”
↩︎ Interpreting log-target predictions and effects“Returns e to the power of expr.”
↩︎ Checkpoint - 5.
“Computes the exponential of the given value minus one.”
↩︎ Undoing the transform: match the inverse to the forward function“Computes the exponential of the given value minus one.”
↩︎ Exam trap 3 - 6.
“predictions$price - predictions$prediction”
↩︎ Compute evaluation metrics after exponentiating, not before - 7.
“Computes the exponential of the given value.”
↩︎ Interpreting log-target predictions and effects
Also cited
- https://scikit-learn.org/stable/modules/model_evaluation.htmlSecondary source
“measuring the distance between predictions y_pred and the true target functional”
↩︎ Exam trap 2