Mitigating Overgeneralization in RND via Spectral Target Design
Abstract
Random Network Distillation (RND) is a scalable novelty signal for reinforcement learning, but it can overgeneralize: the predictor extrapolates the random target off the data manifold, causing novelty scores to collapse on unfamiliar inputs. This paper analyzes this failure through a kernel-theoretic lens. Under a GP target model and an NTK-regime predictor, we show that the expected RND energy is a cross-kernel residual governed jointly by the target covariance kernel and the predictor interpolation kernel. This identifies overgeneralization as spectral under-excitation of modes where predictor residuals would otherwise survive. We then explicitly formulate spectral target design as a principle for shaping RND residual geometry; one concrete instance is replacing to replace the implicit random-network target with a bandwidth-controlled Random Fourier Feature (RFF) target. Numerical experiments across offline D4RL and online Atari benchmarks show this target-design approach improves novelty discrimination and downstream RL performance.