Course
machine-learning-zoomcamp
Question
Homework 2 Q3: filling the horsepower NAs with 0 and with the mean gives me the same RMSE - how do I tell which option is better?
Answer
They are not actually identical, the difference just hides behind rounding. In the pinned 2026
release the validation RMSEs are:
round(rmse_zero, 3) # 2.205
round(rmse_mean, 3) # 2.202 <- better
Round to 3 decimals as the homework asks (the 2026 wording says this explicitly because 2 decimals
turned it into a tie). So the answer is "With mean".
Two things that silently break this question:
-
The mean must be computed on the training split only, never on the full dataset - otherwise
you leak information from validation/test:
hp_mean = df_train.horsepower.mean() # correct
hp_mean = df.horsepower.mean() # leakage
-
Use the same fill value on train, validation and test. Don't recompute the mean per split.
def prepare_X(d, fill_value):
return d[['engine_displacement', 'horsepower', 'vehicle_weight', 'model_year']].fillna(fill_value).values
Checklist
Course
machine-learning-zoomcamp
Question
Homework 2 Q3: filling the horsepower NAs with 0 and with the mean gives me the same RMSE - how do I tell which option is better?
Answer
They are not actually identical, the difference just hides behind rounding. In the pinned 2026
release the validation RMSEs are:
Round to 3 decimals as the homework asks (the 2026 wording says this explicitly because 2 decimals
turned it into a tie). So the answer is "With mean".
Two things that silently break this question:
The mean must be computed on the training split only, never on the full dataset - otherwise
you leak information from validation/test:
Use the same fill value on train, validation and test. Don't recompute the mean per split.
Checklist