ExToken: Structured Exploration for Efficient Vision-Language-Action Reinforcement Fine-tuning
Abstract
Reinforcement Learning (RL) has demonstrated significant potential for improving Vision-Language-Action (VLA) models on complex manipulation tasks. However, its practical scalability remains severely limited by the substantial cost of environmental interactions. In this work, we first investigate the exploration stagnation bottleneck in current VLA-RL frameworks and find that rollouts become increasingly homogeneous as RL training progresses. We further find that a diversity-preserving subset with only half the rollouts retains comparable learning utility, revealing that trajectory diversity can matter more to effective policy learning than the sheer quantity of the explored rollouts. Motivated by these insights, we introduce RL Exploration Token (ExToken), a simple yet general framework that encodes behavioral variation observed in offline demonstration trajectories into discrete exploration conditions. By conditioning the policy on different tokens during rollout collection, ExToken elicits diverse trajectory-level behavioral patterns from the policy, substantially improving exploration efficiency. ExToken further introduces a state-conditioned token selector that routes exploration across tokens during training, adapts its routing strategy online through the RL training, and deterministically selects the optimal token at deployment. Extensive experiments across simulated and real-world robotic manipulation tasks demonstrate that ExToken consistently accelerates convergence, improves task performance, and exhibits strong robustness under highly constrained interaction budgets.