Does an Expert Teacher Help PPO? A Controlled Comparison of Training-Time Guidance Mechanisms on a Reactor Control Benchmark
Nitin Jotwani
Abstract
Facilities that consider reinforcement learning for control almost always possess a trusted classical controller, and a recurring proposal is to keep that controller in the training loop as a *teacher*. Many teacher designs exist, but no controlled experiment compares them: published evidence comes from different learners, tasks, and tuning budgets. We provide that experiment. On a nuclear-microreactor load-following benchmark, one on-policy learner (PPO) is trained under six guidance mechanisms and an unguided control, with a one-step nonlinear model-predictive controller (NMPC) as teacher and classical baseline. Every configuration receives its own hyperparameter search, then twenty evaluation seeds disjoint from the tuning seed, with best-checkpoint selection. Peak accuracy does not separate the methods; reliability separates them sharply, in an order given by one design property: how much of the teacher's behavior enters the learner's training data. Mechanisms that inject teacher actions into the on-policy gradient succeed on 0/20 (action blending) and 6/20 (warmstart) seeds, at or below the unguided control's 8/20, while mechanisms that confine the teacher to shaping states, rewards, or a base action succeed on 19/20 to 20/20 ($p \le 4.5\times10^{-4}$, Fisher exact, Holm-corrected). Guidance reduces the training-time constraint-violation rate by up to $17\times$ and eliminates overpower events; a residual configuration trains to teacher-level accuracy with *zero* violations in 200,600 episodes. Deployment-robustness ablations complete the accounting: under plant–model mismatch the learned policies degrade gracefully while the nominal-model teacher incurs errors one to two orders of magnitude larger; under sensor noise the relationship reverses; under process noise the two are comparable. Resampling our own seeds shows a 3-seed version of this study could have reported any unguided success rate from 0% to 100%, so small-seed comparisons of guidance methods are uninformative.
Chat is not available.
Successful Page Load