Diagnosing Recurrent Computation Reuse under Multitask Optimization
Abstract
Neural-network training can admit multiple solutions that achieve essentially the same loss on the training distribution while differing in their internal organization. This makes learning not only a problem of loss minimization, but also one of solution selection among equally fitting realizations. We study this problem in multitask recurrent neural networks. Prior empirical and theoretical work shows that such models can develop shared features and internal representations, suggesting that multitask training can influence which internal solutions are selected. However, it remains unclear at what level such sharing emerges when training selects among multiple equally fitting solutions. We make this distinction explicit by separating shared state information from shared recurrent dynamics, and study both in controlled sequence tasks with an exactly specified common finite-state transition. We also analyze a restricted optimization construction showing that output loss alone need not favor shared recurrent computation, an issue especially relevant to algorithm learning from input–output supervision.