Attention Geometry Helps Early: Mirror-Descent Updates for Transformer Query and Key Matrices
Abstract
Each row of a self-attention score matrix parameterizes a categorical distribution, yet standard optimizers treat query and key matrices like generic hidden matrices. This raises a natural question: should their updates instead inherit the KL geometry of attention? To answer it, we derive an ideal mirror-descent step in score space, restrict it to changes realizable by shared query and key matrices, and prove a finite-step descent certificate under smoothness assumptions. A damped Fisher/Gauss--Newton approximation makes the target practical with three matrix-free conjugate-gradient iterations. Experiments on FineWeb10B show that the gains concentrate early. Accordingly, we use the spectrally stabilized Q/K update early, while Newton--Muon, a strong matrix optimizer, handles the remaining eligible matrices; we then smoothly hand Q/K off to Newton--Muon. At 124M parameters, this hybrid substantially outperforms Muon. Against Newton--Muon, it improves intermediate-budget performance at both 124M and 353M under matched training schedules. At 353M, it reduces measured training time to intermediate validation targets by up to about 16\%. Longer-horizon runs and ablations show that the gains concentrate early, motivating attention-aware mirror geometry as a promising primitive for frontier optimizers.