State Copying Crowds Out Reasoning: Mechanistic Evidence for Delta Planning in Autoregressive Models
Abstract
Autoregressive models are often used as sequential decision makers by asking them to rewrite the full intermediate state before producing the next action. We study a failure mode of this interface: exact high-entropy state copying can interfere with goal-directed computation. We formulate this as an operational copy--reason interference hypothesis rather than as a directly measurable capacity law. Matched controls separate copying from length: on controlled state-tracking tasks, Full-State Generation degrades much more sharply than an Iso-Length Control and similarly to a Complex-Copy Control. Activation patching on Llama-3.1 8B and 70B recovers target-action evidence when clean residual activations are inserted into late layers of full-state runs, suggesting that goal features remain recoverable but are poorly routed under the copy-heavy interface. Delta-style alternatives reduce this burden. At the 2,000-token Formal Math anchor, Scratchpad Residual improves accuracy from 41.2--68.5\% under Full-State Generation to 97.1--99.4\%; on SWE-bench Lite, it improves pass rates from 12.4--26.8\% to 18.5--34.8\%. The gain has a boundary: probe and task accuracy both decay with long dependency distance. These results argue for separating state maintenance from action generation in long-horizon autoregressive systems.