Auditing Closed-Loop Learning in Recurrent Neural Networks: Reproduction, Robustness, and Generalization
Abstract
Recurrent neural networks are often used as mechanistic models of learning and control, but closed-loop training creates reproducibility challenges because a model's actions alter future inputs. We conduct a claim-level reproducibility study of Ger and Barak's closed-loop RNN learning dynamics, testing independent implementation, seed variation, protocol perturbations, coupled-system diagnostics, and architecture/task transfer. Under a main-text-aligned double-integrator protocol, the trajectory-level peak, not a persistent final gap, reproduces strongly: 50/50 paired seeds show the post-initial open-loop deployed-loss peak, with a mean peak/initial ratio of 19.1, while final open-loop and closed-loop losses converge after open-loop recovery. The spectral stage and coupled-stability diagnostics also reproduce in 50/50 seeds. A targeted A1 analysis separates stability and behavioral tradeoffs: short-horizon improvements coincide with coupled-radius increases in 120/120 runs; long-horizon loss worsening occurs in 62/120, but never without the radius signal. Generalization is hierarchical: GRU variants preserve final-loss divergence, low-rank variants often produce open-loop deployed rollout blow-ups, tanh RNNs preserve the peak signature without a final gap, tracking transfer is strong, and path-integration transfer is weak. These results motivate reporting practices for closed-loop RNN studies: deployed closed-loop loss, peak signatures, paired seeds, spectral stage criteria, coupled-system spectra, feedback strength, rollout horizon, and failure rates.