Scalable Bayesian Posterior Sampling with Weight-Space Symmetries
Abstract
Maximum a posteriori networks are systematically overconfident. The standard post-hoc solution is the Laplace approximation, which fits a Gaussian over the weights whose covariance is the inverse loss curvature at the mode. Common scalable forms of Laplace either discard most covariance between parameters or retain it at a high computational cost. We introduce SymLA, a Laplace posterior that leverages weight-space symmetry to compute and represent the covariance among all network parameters efficiently. We experimentally analyse how the choice of symmetry group controls the trade-off between posterior richness and computational cost. Furthermore, we scale to a pretrained language model, placing a posterior over 43M parameters of DistilBERT. SymLA is competitive with the strongest sampled Laplace posteriors, and our block-diagonal variant performs just as well at a fraction of the cost.