Mode-Controlled Policy Optimization: A Geometry-Aware Recipe for LLM Post-Training
Zhenyu Sun ⋅ Dylan Hadfield-Menell ⋅ Weiguo Feng ⋅ Xiaohuan Zhou ⋅ Yi Zeng
Abstract
LLM post-training is usually framed as a choice of optimizer, but KL-regularized RL also makes a hidden geometric choice: it fits the learned policy to a reward-tilted target by reverse KL. This places PPO, GRPO, RLOO, REINFORCE++, and related methods at a fixed mode-seeking point in a broader design space. We propose \textbf{Mode-Controlled Policy Optimization (MCPO)}, a geometry-aware recipe that replaces this inherited point with an $\alpha$-divergence dial over mode-seeking versus mode-covering behavior; $\alpha{=}0$ recovers the classical reverse-KL recipe. MCPO admits an off-policy objective, a practical grouped importance-weighted estimator, and a baseline-centered variant for stable training. Across mathematical reasoning, agentic transfer, search-based QA, and safety, no single geometry is uniformly best. Average-case math favors stronger concentration, while pass-based math and Search-QA favor broader support; agentic transfer moves with metric and scale; safety exposes a base-dependent frontier in which the same refusal-only signal can transfer broadly or spill into benign boundary cases depending on the reference model's mode landscape. MCPO's value is therefore not a new universal objective, but a way to expose and select the post-training geometry that RL with fixed reverse-KL geometry silently chooses.
Chat is not available.
Successful Page Load