Moment-Constrained Latent Steering for Flow Policy
Abstract
Fine-tuning continuous-time generative models, such as flow models, via reinforcement learning (RL) is a promising approach for complex continuous control, but directly updating the network weights often leads to catastrophic forgetting of the pre-trained prior. Latent steering method mitigates this by freezing the generative backbone and optimizing a latent policy to generate the initial noise. However, existing methods typically enforce hard support boundaries to prevent out-of-distribution (OOD) shifts in the latent space, inducing a fundamental trade-off between prior preservation and policy expressiveness. To resolve this dilemma, we propose Moment Constraint Parameterization, which regulates the latent distribution through a statistical budget. By decoupling this budget into inter-statistic constraints and intra-statistic flexibility, our method acts as a probabilistic safeguard against OOD shifts while providing the degrees of freedom to retain policy expressiveness. Furthermore, while moment constraints define a flexible search space, effectively optimizing the latent policy within it requires accurate gradient signals. To address the optimization lag and bias caused by proxy latent critics in existing methods, we introduce Adjoint Q-Gradient Propagation. By exploiting the white-box structure of flow models, this technique backpropagates exact analytical gradients from the action-space critic directly to the initial noise. Integrating these two mechanisms, we present Moment Constraint Flow Steering (MCFS). Empirical results on D4RL and OGBench show that MCFS consistently delivers stronger online adaptation performance than prior latent steering and flow-policy baselines, with particularly pronounced gains on challenging long-horizon compositional manipulation tasks.