Reward Is Not a Universal Interface for Generative Reinforcement Learning
Abstract
A reward is not a universal RL interface: it becomes a valid update only through the probability object exposed by the policy branch. Autoregressive post-training works cleanly because token ratios are exact, but parallel discrete policies hide within-step dependence and flow/diffusion policies often lack cheap density ratios. We introduce MindRL, a reward-to-update interface controller that translates a shared reward through branch-native score objects while turning factorization, tractability, drift, and smoothness barriers into budgets over block size, serialization, anchors, clipping, reranking, and branch weights. As a closed-loop controller, MindRL makes the shared reward actionable for AR, parallel discrete, flow/diffusion, and AR+flow policies by selecting branch-native update objects and adapting their control budgets. Across language-side AR/parallel-discrete tests and matched flow/diffusion and AR+flow evaluations, MindRL improves task-quality or reward-risk tradeoffs while preserving each branch's native structure.