Weight space gradient descent induces approximate activity descent, making learning rules testable
Abstract
A central question in neuroscience is whether biological neural networks optimise a loss function by adjusting synaptic weights, and if so, by what rule. Candidate rules are difficult to test because synaptic weights cannot be measured at scale in behaving animals, whereas neural activities can be recorded across learning. This raises the question: if a network performs gradient descent on its weights, what should we see in its activities? Here we show that weight-space gradient descent induces kernel-filtered activity descent; when the kernel is approximately diagonal, each neuron follows its own loss gradient. We validate this prediction in ANNs of moderate width and depth, showing that indeed changes in neural activities are correlated with their loss gradients. The same relationship can and should be tested in biological neural networks with existing experimental techniques.