Shutdownable Agents through Length-Neutral Policy Optimization
Abstract
AI agents are increasingly used to autonomously solve long-horizon tasks. Research suggests that models trained in this way might resist shutdown, increasing loss of control risk. We introduce Length-Neutral Policy Optimization (LNPO), which comes in two variants: LNPO-C and LNPO-CR. We derive gradients for each variant. Both are theorized to increase shutdownability while retaining usefulness. We also introduce gridworld environments that act as analogs of scenarios where agents are instrumentally incentivized to resist shutdown via scheming, sandbagging, avoiding monitoring, or acting differently under monitoring. We compare agents trained with Proximal Policy Optimization (PPO) to agents trained with LNPO-C or LNPO-CR and observe that both variants reduce shutdown resistance by 40-71\%, depending on the behavior. We find no performance degradation from either variant when shutdown is irrelevant. Finally, we scale both LNPO variants to LLM fine-tuning on Qwen2.5-7B-Instruct playing Tetris and OLMo-3-7B solving GSM8k problems, where episodes contain offers to trade task points against shutdown probability. Across the two LLM environments, GRPO chooses the trajectory-length-neutral option less often than LNPO-C and LNPO-CR. All three methods improve task performance, and the learned disposition transfers beyond the trajectory length used in training.