CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 3 · Lesson 41/48

    Exponentiating Log-Transformed Targets Before Metrics and Interpretation

    Identify the need to exponentiate log-transformed variables before calculating evaluation metrics or interpreting predictions

    12 min read
    2.08% of exam
    8 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Recognise that a regression model trained on a log-transformed label returns predictions on the log scale
    • Pick the correct inverse function (exp for ln, expm1 for log1p, pow for log10) to bring predictions back to the original units
    • Explain why RMSE, MAE and similar metrics must be computed on exponentiated predictions against the original label, not on log values
    • Interpret predictions and effects from a log-target model in the original business units

    Key concept

    Prediction scale follows target scale — A model returns predictions in whatever scale its label had during training. If you trained on log(y), every prediction is a log value, so you must apply the inverse (exponentiate) before you compare it with real labels, compute an error metric, or report it to anyone.

    1.A log-target model predicts logs, not prices

    Regression targets such as house prices, revenue, salaries or transaction amounts are usually right-skewed: most values are modest and a few are huge. A common fix is to train on the logarithm of the label instead of the raw value. The log pulls in the long tail, makes the residuals more symmetric, and stops a handful of extreme rows from dominating a squared-error loss. In Databricks SQL and PySpark this usually means ln (or log with a single argument, which is the same thing), applied to the label column before training.

    The catch is simple, but the exam keeps testing it. The model has no idea the column used to be a price. It learned to predict whatever numbers you gave it, so if the label was ln(price), then predict returns ln(price). For a house that sells for about $200,000, the natural log is about 12.2, and 12.2 is what comes out of the model. Nothing in the training loop converts it back for you. The exception is a wrapper built for this job, such as scikit-learn's TransformedTargetRegressor, which applies the inverse automatically.

    Checkpoint 1 of 8· Check yourself

    A linear regression was trained on a label column computed as ln(price). For one test row, model.predict returns 12.2. What is the correct reading of this value?

    Sources123

    2.Undoing the transform: match the inverse to the forward function

    To exponentiate correctly, you need to know exactly which log was applied, because each forward function has its own inverse. The natural log ln(x) is undone by exp(x), which is e raised to the prediction. Many pipelines use log1p(x) instead, which is ln(1 + x). It exists because a plain log fails on zero: Databricks returns NULL when the argument is less than or equal to 0, and labels such as sales or counts often contain zeros. The inverse of log1p is expm1, which is e^x − 1. If you apply plain exp to a log1p prediction, every value comes out one unit too high. That error is easy to miss when values are large and very visible when they are near zero.

    Forward transforms applied to a label and the inverse to apply to predictions (PySpark function names)
    Forward transform on the labelInverse to apply to predictionsWhen you see it
    ln(col) / log(col)exp(col)Strictly positive skewed labels such as prices
    log1p(col)expm1(col)Labels that can be exactly 0, such as counts or sales
    log10(col)pow(10, col)Logarithm in Base 10, used less often in ML pipelines
    log2(col)pow(2, col)Base-2 log; invert by raising 2 to the prediction
    PySpark's expm1 is the inverse of log1p: it returns the exponential of the value minus onepython
    from pyspark.sql import functions as dbf
    df = spark.sql("SELECT id AS value FROM RANGE(5)")
    df.select("*", dbf.expm1(df.value)).show()

    Checkpoint 2 of 8· Match them up

    Match each forward transform applied to the label with the function that brings predictions back to original units

    Tap a term, then the definition that fits it.

    Checkpoint 3 of 8· Fill the gap

    The label was transformed with ln. Which PySpark function completes this sample so that it returns the exponential of the column?

    from pyspark.sql import functions as dbf
    df = spark.sql("SELECT id AS value FROM RANGE(5)")
    df.select("*", dbf. ? (df.value)).show()

    Checkpoint 4 of 8· Exam question

    A data scientist trained a regression model to predict home sale prices by first transforming the target using `np.log1p(price)`, since the raw distribution was heavily right-skewed. The model's predictions come out in this transformed scale. Before reporting RMSE to stakeholders in dollar terms, what should the data scientist do?

    Sources425

    3.Compute evaluation metrics after exponentiating, not before

    Regression metrics such as RMSE and MAE are built from the difference between a prediction and a true label. The Databricks glm tutorial below shows this directly: the error column is price - prediction, and RMSE is the square root of the mean squared error. That number is in dollars only because both columns are in dollars. If the model had been trained on ln(price), its prediction column would hold log values. Subtracting those from raw prices produces meaningless numbers, and subtracting them from logged prices produces an error in log units, which is not what a business stakeholder or an exam question asking for RMSE in the original units wants.

    RMSE computed from prediction errors in the label's own units (Databricks SparkR glm tutorial)r
    errors <- select(predictions, predictions$price, predictions$prediction, alias(predictions$price - predictions$prediction, "error"))
    display(errors)
    
    # Calculate RMSE
    head(select(errors, alias(sqrt(sum(errors$error^2 , na.rm = TRUE) / nrow(errors)), "RMSE")))

    So the order of operations is fixed. Transform the label, train, predict (getting log values), exponentiate the predictions, and only then compute the metric against the original, untransformed labels. Do not take a shortcut by computing RMSE on the log scale and exponentiating the result. RMSE is a nonlinear summary over all rows, so exp(RMSE_log) is not the dollar RMSE. The exponentiation has to happen row by row, on the predictions, before any averaging. scikit-learn's TransformedTargetRegressor builds this into the estimator: it fits on the transformed target and maps predictions back with the inverse transform, so predict and score already work in the original units.

    Checkpoint 5 of 8· Put it in order

    Put the steps for evaluating a log-target regression model in the correct order

    1. 1.Fit the regression model on the transformed target
    2. 2.Generate predictions, which come out on the log scale
    3. 3.Map the predictions back to the original space with the inverse transform (exponentiate)
    4. 4.Apply the log transform to the target y

    Checkpoint 6 of 8· Exam question

    A machine learning pipeline predicts subscription lifetime value (LTV) after applying a log transformation to the training target to reduce skew. A product manager asks the data science team to add the model's raw predictions directly to a customer-facing dashboard showing expected dollar value per customer. What should the team do before displaying these numbers?

    Checkpoint 7 of 8· Exam question

    Two candidate regression models were trained to predict a heavily skewed shipping-cost target after log-transforming it. During evaluation, one engineer reports Model A's RMSE computed directly on the log-transformed predictions, while another reports Model B's RMSE after exponentiating predictions back to the original dollar scale. The team wants to pick the better model using these two RMSE values. What is the problem with this comparison?

    Sources63

    4.Interpreting log-target predictions and effects

    The same rule applies when you explain a model, not just when you score it. A prediction of 12.2 means nothing to a pricing team, while about $199,000 does. Every number that leaves the notebook (a dashboard value, a forecast in a report, a threshold for an alert) should be exponentiated with the inverse that matches the forward transform. In Databricks SQL that is exp, which returns e to the power of its argument, and in PySpark it is dbf.exp (or dbf.expm1 for log1p labels).

    Checkpoint 8 of 8· Check yourself

    A model was trained on ln(price) and a Databricks SQL dashboard must show its predictions in dollars. Which function should be applied to the prediction column?

    Effects change shape as well. With a log-transformed label, a linear model's coefficient no longer means 'adds this many dollars'. It acts multiplicatively on the original scale: a one-unit increase in a feature multiplies the predicted value by exp(coefficient). For small coefficients this is roughly a percentage change of 100 × coefficient. Reading a log-scale coefficient as an additive dollar amount is the interpretation version of the same mistake.

    Sources74

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.An RMSE computed between log predictions and log labels can be reported as the model's error in dollars.Why is that wrong?

      That RMSE is in log units. A dollar RMSE needs predictions mapped back to the original space and compared with the original labels.

      Covered in Compute evaluation metrics after exponentiating, not before

    2. 2.You can compute RMSE on the log scale and then exponentiate the RMSE to get the original-scale error.Why is that wrong?

      The metric must measure the distance between original-scale predictions and original labels, so exponentiate each prediction first and then compute the metric.

      Covered in Compute evaluation metrics after exponentiating, not before

    3. 3.Plain exp correctly inverts any log transform, including log1p.Why is that wrong?

      log1p is ln(1 + x), so its inverse is expm1 (the exponential minus one). Using exp leaves every prediction one unit too high.

      Covered in Undoing the transform: match the inverse to the forward function

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 3.
      “TransformedTargetRegressor deals with transforming the target (i.e. log-transform y).”
      ↩︎ A log-target model predicts logs, not prices
      “The predictions are mapped back to the original space via an inverse transform.”
      ↩︎ Compute evaluation metrics after exponentiating, not before
      “TransformedTargetRegressor deals with transforming the target (i.e. log-transform y).”
      ↩︎ Key concept
      “The predictions are mapped back to the original space via an inverse transform.”
      ↩︎ Exam trap 1
      “TransformedTargetRegressor transforms the targets y before fitting a regression model.”
      ↩︎ Checkpoint
    2. 5.
      “Computes the exponential of the given value minus one.”
      ↩︎ Undoing the transform: match the inverse to the forward function
      “Computes the exponential of the given value minus one.”
      ↩︎ Exam trap 3

    Also cited

    Spotted a mistake, or was something unclear? Tell us.