Recent Training is Protected by Adam's Second Moment
Shrey Biswas
Abstract
Training does not only move weights; it also raises Adam's second moment on certain rows. That excess later blocks new writes from permanently overwriting the older ones by dividing down their updates on those rows, until the protection decays. We investigate by training a model on a task with two equally valid solutions, so ordinary data favours neither. By inserting two lessons - periods where only one solution is correct - we analyse how the model 'writes' these lessons about a specific solution. We first apply one lesson so that the model learns that solution. We see that waiting 300 steps before applying the opposing lesson results in the network drifting back to the first solution (first-lesson preference climbing from near zero after the second lesson, to 0.40-0.97), while waiting 3,000 steps results in the second lesson overwriting the first (first-lesson preference stays near zero, 0.01-0.10). We then find that this protection decays within 8% of the excess second moment's $-1/\ln\beta_2$ decay; and establish a law predicting whether the second lesson sticks or not based on the ratio of the two lessons' integrated writes exceeding 0.63. We establish this to be causal; upon clearing the excess second moment, localised on the rows the second lesson is about to write to, the protection vanishes and the second lesson now consistently overwrites the first. This is as effective as fully resetting optimiser state, and more effective than lowering $\beta_2$ - which backfires as predicted by our law, since the second lesson's write becomes insufficient. We conclude that the second moment's excess is a protection mechanism for previous writes, one that is readable from two optimiser-state snapshots, editable in place, and reset by fine-tuning schedules as if it held nothing.
Chat is not available.
Successful Page Load