Entropy Regularization: A Free Correction to Cross-Entropy for Verifiable Tasks
Abstract
Large language models are often post-trained on expert demonstrations using cross-entropy (CE), even when the downstream objective is not to imitate the demonstrated solution but to produce any output accepted by a verifier. This mismatch is particularly salient in verifiable domains with multiple correct solutions, such as mathematical reasoning and code generation, where training data may contain only one expert solution per problem. We show that minimizing cross-entropy can be fundamentally misaligned with minimizing verifier risk; two policies can assign identical likelihood to the observed demonstrations while placing substantially different probability mass on incorrect outputs. We formalize this phenomenon through a learning-theoretic counterexample in which CE minimization selects a suboptimal policy. Our analysis suggests that controlling the support of the learned policy can resolve this ambiguity by discouraging probability mass from spreading to unsupported outputs. Since support size is non-differentiable and computationally intractable, we propose \textbf{entropy-regularized cross-entropy (ER-CE)}, using token-level Shannon entropy as a tractable proxy. We establish conditions under which ER-CE improves upon standard CE and requires no additional samples beyond ordinary supervised fine-tuning. Finally, across mathematical reasoning and code-generation benchmarks, we find that entropy-regularized training consistently improves verifier accuracy over standard cross-entropy. Together, our results identify a simple failure mode of imitation-based post-training in verifiable tasks and provide a practical objective that better aligns learning from demonstrations with the goal of producing correct outputs.