3.13 Exponentiating Log-Transformed Predictions Before Evaluation

Key Takeaways

  • A model trained on z=ln⁡(y+1)z = \ln(y+1) predicts on the log scale; its raw predictions are not in the target's units.

  • Metrics computed on z^\hat{z} measure logarithmic error — RMSE on the log scale is effectively RMSLE, not dollars.

  • Invert with y^=exp⁡(z^)−1\hat{y} = \exp(\hat{z}) - 1 (np.expm1, F.expm1) before computing business metrics or reporting predictions.

  • Use the exact inverse of the transform applied: expm1 undoes log1p, exp undoes log, and 10z^10^{\hat{z}} undoes log⁡10\log_{10}.

  • Comparing a log-target model against a raw-target model is only valid after both are evaluated in the same original units.

Last updated: August 2026

3.13 Exponentiating Log-Transformed Predictions Before Evaluation

When a skewed target is log-transformed for training (Section 2.5), everything the model emits lives on the log scale. Forgetting to come back is one of the most common silent errors in applied regression, and the exam names it explicitly.


Evaluating Models Trained on Log-Transformed Targets

When predicting highly skewed targets (e.g. house prices, customer spend, income), practitioners frequently apply a logarithmic transformation during data preparation:

z=ln⁡(y+1)(or z=log⁡10(y))z = \ln(y + 1) \quad (\text{or } z = \log_{10}(y))
+-----------------------------------------------------------------------------+
|                LOG-TRANSFORMED TARGET EVALUATION WORKFLOW                   |
|                                                                             |
|   Training Phase:                                                           |
|   Target y ---> [ z = ln(y + 1) ] ---> Model Fits on z ---> Predicts z_hat  |
|                                                                             |
|   INCORRECT EVALUATION (Log-Scale Error):                                   |
|   RMSE(z, z_hat) = 0.35  <-- Reflects Root Mean Squared Log Error (RMSLE)   |
|                              DO NOT compare directly to un-logged RMSE!     |
|                                                                             |
|   CORRECT EVALUATION (Original Dollar/Unit Scale):                          |
|   z_hat ---> [ y_hat = exp(z_hat) - 1 ] ---> RMSE(y, y_hat) = $42,500       |
|   * Allows direct apples-to-apples comparison with baseline dollar models!  |
+-----------------------------------------------------------------------------+

Important

Exam Rule for Log Transformations: Computing metrics on predicted log values z^\hat{z} measures error on the logarithmic scale (equivalent to RMSLE). To evaluate real-world business impact and compare fairly against models trained on the raw target, you must apply the inverse transformation y^=exp⁡(z^)−1\hat{y} = \exp(\hat{z}) - 1 (for log1p) and calculate metrics between true yy and y^\hat{y}.


Matching the Inverse to the Transform

The inverse must exactly undo what was applied. Mixing them produces predictions that are wrong by a constant factor, which is easy to miss because they still look plausible.

Applied during preparationInverse before evaluationNumPyPySpark
z=ln⁡(y)z = \ln(y)y^=ez^\hat{y} = e^{\hat{z}}np.expF.exp
z=ln⁡(y+1)z = \ln(y+1) (log1p)y^=ez^−1\hat{y} = e^{\hat{z}} - 1np.expm1F.expm1
z=log⁡10(y)z = \log_{10}(y)y^=10z^\hat{y} = 10^{\hat{z}}np.power(10, z)F.pow(F.lit(10), z)
z=yz = \sqrt{y}y^=z^2\hat{y} = \hat{z}^2np.squareF.pow(z, 2)

Using np.exp where np.expm1 was required shifts every prediction by exactly 1 unit — negligible on a $500,000 house, material on a 3-unit demand forecast.

Two Distinct Failures

  1. Reporting the wrong number. "The model's RMSE is 0.28" is meaningless to a stakeholder expecting dollars. On a natural-log target it describes a multiplicative error: a typical prediction is off by a factor of e0.28≈1.32e^{0.28} \approx 1.32, so about 32% high or low. The shorthand "an RMSE of rr on a log target is about an r×100%r \times 100\% relative error" is only a small-rr approximation, and it understates the true figure as rr grows — 0.28 reads as 28% but is really 32%.
  2. Comparing incomparable models. A log-target model showing RMSE 0.28 and a raw-target model showing RMSE 42,000 cannot be ranked against each other. Invert the first, then compute both RMSEs in dollars.

A note on the mean. Because exp⁡\exp is convex, exp⁡(E[z])≤E[exp⁡(z)]\exp(\mathbb{E}[z]) \le \mathbb{E}[\exp(z)] (Jensen's inequality), so naively exponentiating a prediction of the log-mean slightly under-predicts the mean on the original scale. The exam expects the straightforward expm1 inversion; the bias correction matters mainly when the residual variance is large.

A Worked Number

Take a house whose true price is $400,000, on a log1p target.

  1. Wrong — report the log-scale residual. The true value on the log scale is ln⁡(400,001)≈12.899\ln(400{,}001) \approx 12.899. A prediction of z^=12.9\hat{z} = 12.9 gives a log-scale error of about 0.0010.001: a number that sounds like near-perfection and means nothing to a stakeholder.
  2. Right — invert first. y^=e12.9−1≈400,311\hat{y} = e^{12.9} - 1 \approx 400{,}311, so the dollar-scale error is roughly $311.

Now move the prediction to z^=13.2\hat{z} = 13.2. The log-scale error grows to about 0.300.30, which still reads as small, but e13.2−1≈540,364e^{13.2} - 1 \approx 540{,}364 — a miss of roughly $140,000. That is the property which makes log-scale metrics deceptive: a constant error on the log scale is a constant proportional error on the original scale, so the same 0.300.30 is $311-sized on a cheap house and six figures on an expensive one.

Why log1p Rather Than Plain log

log1p exists for the zero problem. A target with legitimate zeros — units sold, claim amounts, session minutes — cannot be passed to ln⁡\ln, which is undefined at 0. Shifting by one maps ln⁡(0+1)=0\ln(0 + 1) = 0, keeping those rows in the data set at a finite value. The cost is that the shift must be remembered on the way back: expm1, never exp.

Interpreting Predictions, Not Only Metrics

The objective names interpretation alongside metric calculation, and the trap runs in both directions. On a log target an effect is multiplicative, not additive: a prediction that rises by 0.100.10 on the log scale corresponds to roughly a 10.5% increase in the target (e0.10−1=0.105e^{0.10} - 1 = 0.105), not a 0.10-unit increase. A prediction interval has to be exponentiated endpoint by endpoint, which is why a correctly inverted interval is asymmetric in dollars around the inverted point estimate.


Worked Implementation: Inverting a Log Target in PySpark

PySpark ML: Multi-Metric Evaluation & Log-Target Inversion

from pyspark.sql import functions as F
from pyspark.ml.evaluation import RegressionEvaluator

# Load predictions DataFrame containing log-transformed target
# actual: 'log_price', predicted: 'prediction' (log scale)
scored_df = spark.table("lakehouse_gold.real_estate_predictions")

# 1. Evaluate metrics on Logarithmic Scale (RMSLE)
evaluator_rmse = RegressionEvaluator(labelCol="log_price", predictionCol="prediction", metricName="rmse")
evaluator_r2 = RegressionEvaluator(labelCol="log_price", predictionCol="prediction", metricName="r2")

log_rmse = evaluator_rmse.evaluate(scored_df)
log_r2 = evaluator_r2.evaluate(scored_df)
print(f"Log-Scale RMSE: {log_rmse:.4f}, Log-Scale R2: {log_r2:.4f}")

# 2. Invert predictions back to original dollar scale: y = exp(z) - 1
inverted_df = scored_df.withColumn(
    "actual_price", F.expm1(F.col("log_price"))
).withColumn(
    "predicted_price", F.expm1(F.col("prediction"))
)

# 3. Evaluate metrics on True Original Dollar Scale
evaluator_dollar_rmse = RegressionEvaluator(
    labelCol="actual_price",
    predictionCol="predicted_price",
    metricName="rmse"
)
evaluator_dollar_mae = RegressionEvaluator(
    labelCol="actual_price",
    predictionCol="predicted_price",
    metricName="mae"
)

dollar_rmse = evaluator_dollar_rmse.evaluate(inverted_df)
dollar_mae = evaluator_dollar_mae.evaluate(inverted_df)

print(f"True Dollar Scale RMSE: ${dollar_rmse:,.2f}")
print(f"True Dollar Scale MAE:  ${dollar_mae:,.2f}")
Test Your Knowledge

A data science team trains a gradient boosted regression model to predict housing prices where the target was transformed using z=log⁡(y+1)z = \log(y + 1). The model achieves a validation Root Mean Squared Error of 0.28 on zz. How should the team compute the true dollar-scale RMSE to report to executive stakeholders?

A

Multiply 0.28 by the average home price in the training dataset.

B

Apply the inverse transformation y^=exp⁡(z^)−1\hat{y} = \exp(\hat{z}) - 1 to the predictions, and calculate the RMSE between actual dollar prices yy and predicted dollar prices y^\hat{y}.

C

Compute 100.2810^{0.28} and add the training set mean residual.

D

Report 0.28 directly because RMSE is invariant to logarithmic transformations.

Test Your Knowledge

A model is trained on y_log = np.log1p(y) and predicts 12.5 for a house. Which conversion returns the predicted price in dollars?

A

12.5 ** 2

B

np.exp(12.5)

C

np.expm1(12.5), which computes exp(12.5) - 1, the exact inverse of log1p

D

12.5 * y.mean()

Test Your Knowledge

A team reports that their log-target price model achieves RMSE 0.31, and compares it favourably against a colleague's raw-target model at RMSE 38,000. Why is this comparison invalid?

A

RMSE cannot be used for regression models trained on transformed targets

B

The 0.31 is measured on the logarithmic scale (effectively RMSLE) rather than in dollars; the predictions must be inverted with expm1 and the RMSE recomputed against actual dollar prices before the two models can be compared

C

The colleague's model must also be log-transformed for the numbers to be comparable

D

RMSE is only comparable when both models use the same algorithm

Sections you finish are checked off in the contents.