Sign Descent Sees the Internal Basis: Opposite Learning Orders in Deep Matrix Factorization
Abstract
Sign descent normalizes the gradient entrywise, so unlike gradient descent and spectral descent it is not invariant under orthogonal changes of the internal bases of a factorized model. What does this do to the order in which a target is learned? We study this question in deep matrix factorization, where the spectral initialization of Saxe et al. reduces gradient and spectral dynamics to scalar equations, but entrywise normalization generally breaks the diagonalization underlying this reduction. Motivated by Garrod et al., we identify Sylvester–Hadamard initializations for which the sign map preserves the spectral structure, so the diagonalization survives and layerwise sign descent reduces to dynamics on the singular values. Our construction lets us choose the internal bases, and we compare two choices that agree on the prediction matrix, on the layer singular values and on the entire gradient- and spectral-flow trajectory. Under sign descent, however, one choice recovers spectral dynamics, while the other couples the singular values through the Hadamard transform. We show that already for a target with two distinct singular values, the two learn the modes in opposite orders.