Centered: Notes on the ubiquitous topic of data normalisation
Crossrail Place Footbridge, October 2020 Data normalisation, transforming the variables in a regression to be mean zero with variance one (sometimes called standardisation), comes up frequently in work so it felt worth writing some notes on it. Variables in regressions are often on different scales (e.g. if we are predicting shops’ revenue using inputs of the local population count in 1000s and the shop’s online rating between 1 and 5) the regression coefficients on the larger scale variables will generally be smaller as they have more variation. A direct comparison of the coefficients will suggest the variable on the smaller scale is more important, when in practice it may vary relatively little e.g online ratings can cluster at certain levels like four. ...