Reinforcing Internal Computations for Language Reasoning
Yibin Wang ⋅ Kaili Zhao ⋅ Jianyu Wang ⋅ Xiangjun Fan ⋅ Chengzhi Mao
Abstract
We find that Mixture-of-Experts (MoE) models contain alternative internal computation paths that enable advanced reasoning. Standard MoE inference uses deterministic routing, which causes not only an over-reliance on the top-$K$ experts at each token, but also collaterally suppresses alternative expert paths that may yield correct solutions. We demonstrate that applying reinforcement learning to the routing policy to explore and reinforce these alternative paths improves reasoning capabilities, providing a stronger mechanism beyond token-level optimization alone. Experiments demonstrate significantly improved absolute accuracy by up to 4% over token-only baselines, across several mathematical reasoning benchmarks. Our results also show favorable transfer on held-out general reasoning tasks outside the training domain. Since our approach samples experts without increasing the active parameter count, it is compatible with the standard MoE computation budget. Our results suggest that reinforcing the internal computation policy can serve as a powerful complement to existing training algorithms that optimize the output-token policy.
Chat is not available.
Successful Page Load