Enhancing Reasoning Reliability in LLMs with Tandem Training: Chess as a Model System
Abstract
Outcome-only reinforcement learning with verifiable rewards (RLVR) can improve final-answer accuracy without ensuring reliable reasoning. Reasoning reliability also concerns whether claims are factual, how confident the model is in correct content, and whether a weaker model can successfully continue its reasoning. Assessing these properties in free-form language is challenging, however, because claims can be ambiguous and task constraints are often implicit. We therefore use chess as a model system, where explicit rules, verifiable board states, and standardized move notation support systematic evaluation. Starting from the same base model, we compare conventional RLVR with tandem training, which co-generates rollouts with a frozen copy of the base model. Specifically, we evaluate reasoning reliability through three optics: factuality, conditional likelihood of correct move notation, and handoff robustness with the weaker base model. Within factuality, we define \emph{terminal-successor hallucination} (TSH) as a trace that presents a continuation from an environment-verified terminal state as actual, even though the environment admits no legal outgoing transition from that state. Tandem training achieves higher factuality (including lower TSH incidence), greater conditional likelihood of correct move notation and the checkmate marker, and stronger handoff performance. These findings support co-generation as a way to improve reasoning reliability beyond final-answer accuracy.