Training LLMs to Follow Reasoning Instructions\\with Minimal Edits
Abstract
Large language models increasingly rely on extended reasoning traces, yet the training signals used to shape this behavior remain coarse. Supervised fine-tuning on full trajectories forces the model to imitate an entire demonstration, overconstraining its output distribution, while trajectory-level reinforcement learning provides only weak credit assignment, rewarding or penalizing whole completions without identifying where they succeed or fail. We propose edit-based training, a general framework for learning from a model's own trajectories via localized textual corrections. Given a sampled completion, an editor identifies the first point of failure under a specified criterion and proposes a minimal replacement span. The model is then trained only on these edited spans, leaving the remainder of the trajectory unsupervised. By design, this changes the model only where its reasoning first breaks down and preserves its existing reasoning patterns everywhere else, providing targeted supervision that avoids the entropy collapse and capability loss caused by imitating full demonstrations. The method applies both offline and online, with the latter continually adapting supervision to the model's evolving failure modes. We instantiate this approach on reasoning instruction following with Qwen3-32B on the ReasonIF benchmark. Compared to full-trajectory imitation and trajectory-level reinforcement learning baselines, edit-based training substantially improves compliance with the reasoning instruction without loss of the model's skills on untargeted tasks. These results suggest that localized correction is an effective and general complement to these approaches for correcting the reasoning of LLMs.