Deconstructing Gradient Reuse: A Prospective Target-Permutation Diagnostic for Online Learning
Abstract
Previous gradients can improve a model prior to the arrival of new labels. However, empirical improvement alone does not reveal whether this utility stems from learning correct input-target pairings or merely from absorbing marginal statistics from the data that survive target reassignment. To resolve this ambiguity, this work introduces a prospective target-permutation test. By comparing matched virtual updates using true versus permuted targets, the test isolates the specific value of correct pairings while leaving the online learner unchanged. The analysis formalizes this distinction through exact decompositions of squared loss and cross-entropy, explicitly separating the assignment-dependent terms neutralized by permutation from the marginal terms retained in expectation. Evaluating this framework across chronological regression, image classification, and sequence modeling reveals distinct learning regimes. The results show that shallow regression improvements largely survive permutation, indicating a reliance on general statistics; conversely, classification and sequence models depend fundamentally on correct pairings. Supported by rigorous diagnostics, including repeated permutations, bootstrap sensitivity, and step analyses, this test transforms the vague assertion that an old gradient "transfers" into a controlled, quantitative measure of the precise information driving model improvement.