Diverse Alternative Reasoning Traces for Hard Math Problems
Abstract
Post-training large language models (LLMs) for mathematical reasoning often relies on traces whose final answers pass a verifier. For hard math problems, however, such traces can be rare under independent Monte Carlo (MC) sampling. To address this, we propose Diverse Alternative Reasoning Traces (DART), a targeted proposal mechanism that starts from an incorrect trace, ranks its individual reasoning steps by position-weighted uncertainty, and prompts the model to replace selected steps with deliberately distinct alternatives before completing the trace. On MATH-500, DART recovers 170 greedy failures compared with 148 for matched step-level MC, a relative increase of 15%. On subsets of DeepMath-103K and OpenMathInstruct-2, it solves up to twice as many model-hard problems as MC sampling at a matched number of completed traces. This comes at 39% more generation FLOPs per completed trace than MC sampling, but 30% fewer generation FLOPs per correct trace. Post-training Llama-3.1-8B-Instruct on the correct traces obtained by DART outperforms all compared baselines in five of six evaluations. More broadly, our method provides a practical way to expand exploration beyond conventional sampling to solve more model-hard problems and obtain additional correct traces for post-training.