MTMA: Multi-Turn RL for Long-Horizon Moral Alignment
Abstract
Current RL alignment methods optimize for narrow safety properties (calibrated refusals, jailbreak robustness) on single-turn chat interactions, while models are increasingly deployed to act over long multi-turn horizons. We hypothesize that learning from experience over multi-turn trajectories aligns models more deeply, i.e. more robust and generalizable than single-turn training. We introduce Multi-Turn Moral Alignment (MTMA), a critic-free group-based RL algorithm that uses dense per-turn deontological rewards (penalizing deception, violence, etc.) to train models to act ethically throughout long-horizon trajectories (64-128 turns, 40-80K tokens). We fine-tune Qwen3.5-{2B,4B,9B} models in MACHIAVELLI with MTMA and reduce ethical violations per turn by up to 25% on trained games and 17% on held-out games at twice the training horizon, outperforming existing multi-turn RL methods. On out-of-distribution AIRiskDilemmas evaluations, MTMA models choose the risk-avoidant action up to 8 points more often without any explicit AI-assistant fine-tuning. MTMA also improves multi-turn normative robustness: MTMA models are about 25% more robust to presentation order and about 30% less swayed by user pressure.