Curriculum Learning with Internal-State-Aware Reward Allocation for Animal Behavior Shaping
Abstract
Training animals to perform complex behavioral tasks is a major bottleneck in systems neuroscience. A common training strategy is shaping: animals begin with easier versions of a task and progress toward the target behavior as their performance improves. This process resembles automated curriculum learning (ACL), where task difficulty is adapted to the learner's performance. Standard ACL formulations, however, typically assume that the learner remains available for continued training. In animal shaping, participation is not guaranteed. Animals may disengage after repeated failures or effortful trials. Reward can help sustain participation, but its motivational value declines with satiety and each session has a finite reward budget. Here, we formulate shaping as a sequential control problem over task difficulty and reward magnitude. Difficulty determines what the animal practices, whereas reward allocation influences whether future training opportunities remain available. We characterize how these decisions should be coordinated using a tractable three-step simulation, then develop a scalable strategy that adapts difficulty from recent performance and reward from the animal's within-session state. Adaptive reward allocation approaches a full-state dynamic-programming benchmark and provides the greatest benefit when declining participation would otherwise shorten training sessions. The same principle reduces sessions to mastery in longer sequential and perceptual-learning tasks. These results establish limited trial availability as a key feature of animal shaping and suggest that efficient training requires managing not only what an animal practices, but also how limited reward is allocated to sustain future learning opportunities.