Local Learning Coefficients Reveal How Data Complexity Shapes Solution Degeneracy, Independently of Architecture
Abstract
Understanding generalization in deep neural networks requires understanding the interaction between architecture and data. We study this interaction using the local learning coefficient (LLC), which measures the degeneracy of the loss landscape around a learned solution. We train three architectures on datasets whose class-conditional moments match those of CIFAR-10 up to a given order, and estimate the LLC of the resulting solutions. The LLC increases with the moment order for all three architectures: data retaining higher-order structure yield less degenerate solutions. The absolute LLC values differ substantially across architectures, but normalizing by each architecture's real-data LLC makes the three trajectories nearly coincide. This suggests that the architecture determines the overall size of the LLC, while how the LLC grows with the moment order is nearly the same across architectures. The complexity of the data thus affects the learned solution through the degeneracy structure of the loss landscape.