SDAE: Semantic-Diversity-Aware Exploration for Efficient Reinforcement Learning in Large Language Models
Abstract
Reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models. However, as training progresses, models often converge to narrow solution paths, leading to diversity collapse. Existing methods alleviate this phenomenon through token-level entropy regularization or negative sample reinforcement, yet these works have treated responses as atomic units distinguished only by binary correctness labels, overlooking substantial semantic heterogeneity among responses under the same label—among incorrect responses, the degree and type vary significantly; among correct responses, both conventional and novel strategies exist. Uniformly reinforcing all correct responses leads to policy convergence, while uniformly penalizing all incorrect ones may suppress promising directions. Building on this insight, we propose Semantic Diversity-Aware Exploration (SDAE), which leverages geometric structures in semantic space to guide exploration. SDAE computes the semantic centroid of positive samples and modulates advantage functions based on each response's distance to this centroid, encouraging novel correctness, penalizing divergent errors, while preserving improvement potential in locally flawed responses. Experiments demonstrate that SDAE consistently outperforms strong baselines across competition-level reasoning benchmarks, achieving a 13.3% improvement in Pass@256 on AIME25. Our code is available at https://anonymous.4open.science/r/SDAEcode-3553.