Same Training Loss, Different Recovery: Optimizer-Dependent Solution Selection in Matrix Sensing
Abstract
In an overparameterized problem, exact minimization of the training objective does not guarantee recovery of the ground truth solution. A simple non-convex example is low-rank matrix sensing, where the solution can be planted, and recovery can be measured directly. We compare SGD, Global RMS, RMSProp, SignSGD, and Muon using matched initializations and minibatches. All five methods reach almost zero training loss, but the matrix solutions they find are very different. SGD reaches a high-effective-rank interpolating solution early, whereas RMSProp and Muon reduce the training loss more slowly and recover the target by the final checkpoint. To compare the solutions found by the different methods, we follow the size and direction of the updates and rotate the whole problem. Global RMS keeps the update parallel to the gradient but does not recover, while momentum-free SignSGD changes its direction but also does not recover. The rotated runs further show that basis dependence is not what separates the methods. In this experiment, none of these properties predicts recovery on its own. We still do not know why RMSProp and Muon select the planted solution in this experiment.