Grid Cells as Spectral Successor Features: A Reusable Basis for Continual Reinforcement Learning
Abstract
A reinforcement-learning agent moving between tasks in a fixed environment faces a mismatch: the reward changes, but the geometry does not. Methods that learn a representation and a value function jointly relearn both at every task switch, which is slow and prone to forgetting. We take the identification of grid cells with eigenvectors of the successor representation as a design principle, and build the feature basis from the low-frequency eigenvectors of the symmetrised random-walk graph Laplacian, defined over state-action pairs rather than states so that the features encode what an action does. Because this basis diagonalises the transition operator, the successor-feature transform is exactly diagonal with closed-form entries fixed by the Laplacian eigenvalues and the discount. No temporal-difference learning is needed for it, and neither a successor-feature network nor a replay buffer is required: the only learned quantity is a linear reward map, re-estimated per task. On two goal locations in a two-room maze sharing an identical transition structure, the resulting agent holds near-ceiling return across every task switch, where a successor-feature baseline that learns its features collapses at each change. The cost is a slightly lower ceiling imposed by the linear parameterisation. The basis here is computed from the true transition graph, so this arm is an oracle upper bound rather than a method competing on equal information; learning the basis online is in progress.