Stability in Multi-Step Reasoning via Jacobian-based Error Accumulation Analysis
Abstract
We study the accumulation of test errors in language models performing multi-step reasoning. Although longer reasoning improves model capabilities by scaling test-time computation, test error compounds in autoregressive generation and can grow substantially across steps. The key factors determining the stability of such reasoning remain unclear. To address this question, we analyze the test error of multi-step reasoning under supervised fine-tuning. We first present a test error bound governed by the product of input Jacobian spectral norms across generation steps. This product, summarized as an error amplification factor, scales exponentially with steps and controls the reasoning stability. Then, we explicitly analyze this factor in transformers trained to predict linear and quadratic function weights. We prove that the transformer converges to a solution where this amplification factor decays, yielding near-zero test loss even over longer steps. Based on the analysis, we propose a training method to suppress the error amplification factor by combining (i) chain-of-thought length compression that reduces the reasoning steps, and (ii) quantization-aware training that provably regularizes the input Jacobian norms. We validate our method by fine-tuning language models on reasoning tasks that require executing graph algorithms and tracking variable states through logical operations. Across seven tasks, our method improves over baselines by 3.5\% on average, and by 8.2\% in length generalization evaluations. We further validate that it reduces the error amplification factor in fine-tuned models.