What the Geometry of Good Models Tells Us
Abstract
Practitioners using machine learning on noisy tabular data can often find simpler, more interpretable models with the same accuracy as black box models. In some cases, this is because the noise in the data induces a large Rashomon set, that is, a large set of near-optimal models, which then frequently contains simpler models. However, our understanding of this effect from discrete hypothesis classes --- the Rashomon set's size, the diversity of its predictions, and implicit regularization --- leads to a contradiction in continuous hypothesis spaces, and thus does not directly establish whether noise should necessarily induce simpler or more diverse models. We reconcile the seemingly conflicting effects by showing geometrically that noise drives simplicity through separate mechanisms for linear, logistic, and generalized additive models. We introduce the Rashomon reducibility cost to quantify whether a given model simplification is possible within the Rashomon set, and characterize how this cost changes under general data properties such as label noise. For practitioners working with noisy tabular data, our results suggest that simpler, more interpretable models will be included in the Rashomon set, so that black box models will generally not be needed.