Multiplicative vs. Additive State Coupling in Learned Optimizer Schedules
M. Harper Langston ⋅ Pierre-David Letourneau ⋅ Dalton Jones ⋅ Roland Memisevic ⋅ Peter Scott ⋅ Mingu Lee ⋅ Harris Teague ⋅ Richard A Lethin
Abstract
We learn schedules for a continuous penalty relaxation of MAX-CUT by backpropagating through an unrolled classical solver. A state-blind schedule generalizes to instances substantially larger than those used for training. A state-adaptive bilinear schedule also generalizes out of distribution, outperforming its untrained baseline across 15 seeds ($t=4.1$). By contrast, an additive ablation with identical inputs and output heads is statistically indistinguishable from the bilinear model in distribution but performs worse than its own baseline out of distribution ($t=-11.7$). This separation appears without regularization and persists across a $30$-round weight-decay sweep: the bilinear advantage remains significant at every tested value, whereas the additive model never recovers. The results in this setting isolate how solver state enters the schedule, rather than state adaptivity alone, as the architectural distinction associated with out-of-distribution transfer.
Chat is not available.
Successful Page Load