Causal Intervention on Agent Reasoning: Not All Reasoning Is Load-Bearing
Abstract
Agent reasoning traces are often treated as explanations of why an agent acted, but agreement between reasoning and action is correlational: the same action can follow reasoning that caused it, or reasoning that merely accompanied a decision made elsewhere. We introduce checkpoint resumption, a method for causal intervention in long-horizon agent trajectories that reconstructs a stored checkpoint, substitutes its private reasoning, and samples only the next action without re-executing the preceding trajectory. From a corpus of 324 trajectories from Agents' Last Exam, we find that reasoning is not uniformly load-bearing. Corrupting a non-recoverable derived value produces no detectable change in the next command beyond a meaning-preserving paraphrase control, with a corruption effect of −2.9 percentage points on Opus (95% CI [−14.5, +8.8]) and 0.0 on GPT. In contrast, replacing a stated intention changes the next action in 90% of Opus cases and 80% of GPT cases. A second crossed intervention shows that obedience depends on how an instruction is written rather than what it asks for: on Opus the same instruction is obeyed 87.5% of the time when phrased as the model's own reasoning against 20.0% as an external directive, with no main effect of task consistency. These findings replicate across two frontier models, two agent harnesses, and two providers. Observational reasoning and action consistency can therefore conceal substantially different causal roles within the same trace, so verifying an agent from its trace requires intervention rather than inspection.