VAM: Verbalized Action Masking for Controllable Exploration with Verifiable Rewards
Abstract
Exploration remains a key bottleneck in reinforcement learning with verifiable rewards for large language models, where repeated actions and identical rewards can leave many sampled training positions with little comparative signal for learning. We propose Verbalized Action Masking, or VAM, a simple mechanism for improving within-state exploration by verbalizing an allowed action set in the prompt and iteratively pruning previously sampled in-mask actions when the target is missed. This directs subsequent rollout groups toward previously untried alternatives while keeping each group on-policy for its mask-conditioned prompt. We evaluate VAM in chess under both fixed-dataset and engine-play training regimes. VAM maintains a substantially larger prompt-level effective batch size than Chess-R1, indicating that more sampled positions produce nonconstant verifier rewards and contribute informative relative-reward signal to each update. Across Qwen2.5 and Qwen3 models, VAM achieves better or comparable held-out puzzle accuracy and consistently lower full-game average centipawn loss. These empirical results show that VAM improves exploration efficiency and makes better use of RLVR training resources by replacing redundant action sampling with informative alternatives.