Grid Cells as Spectral Successor Features: A Reusable Basis for Continual Reinforcement Learning
Élodie Côté-Gauthier ⋅ Isabeau Prémont-Schwarz ⋅ Doina Precup
Abstract
Continual reinforcement learning requires agents to adapt when tasks change while preserving reusable structure. Successor features separate task-specific reward from task-independent dynamics, but depend strongly on the feature basis. Taking the identification of grid cells with eigenvectors of the successor representation as a design principle, we build that basis from the low-frequency eigenvectors of the random-walk graph Laplacian. Because this basis diagonalises the transition operator, the successor-feature transform becomes exactly diagonal with closed-form entries $\theta_i = 1/(1-\gamma(1-\mu_i))$, removing an entire learning problem: no successor-feature network, no replay buffer, and a single linear regression per task. We further show the basis must be built over state-action pairs, since invariance under the \emph{average} transition operator does not give the per-action invariance an action value requires. On a two-goal gridworld the agent holds near-ceiling return across every task switch, where a baseline that learns its features from the same pixels collapses at each change; the cost is a slightly lower ceiling imposed by the linear parameterisation.
Chat is not available.
Successful Page Load