Diagnosing and Reducing the Reasoning-to-Execution Gap in LLM Sequential Decision Agents
Abstract
LLM agents in sequential decision settings often fail at doing, not knowing: they can state a correct belief and still take an action that belief does not justify. Harness-reliability metrics cannot see this, because they check whether an action was correct, not whether it followed from the agent's own stated reasoning, and realistic settings have no ground-truth optimal action to check against. We build a minimal repeated-pricing game small enough to admit an exact expected-value oracle, paired with same-pass belief elicitation, giving a two-oracle decomposition of failure into estimation error (is the stated belief close to the correct posterior?) and execution error (does the chosen action follow from the agent's own stated belief?). On a weak model both channels are severely broken (directional agreement 0.524, barely above chance; mean EV-realization ratio -0.369). Capability relocates the gap rather than closing it: the frontier model's beliefs are largely correct (TV distance 0.063) yet it still forfeits value on 66% of decisions relative to its own stated belief. We then test a no-training scaffold, a belief-state carry plus a mandated re-evaluation step, against a token-matched control that isolates structure from reasoning budget. It does not close the gap either: on the frontier it widens the estimation axis, worst on the belief-carry component (TV 0.088 vs. 0.052 for the control), which we trace to double-counting of already-summarized evidence and confirm with a signed over-concentration fingerprint replicated on a second frontier model. Replacing raw history with its lossless sufficient statistic repairs the belief axis at negative token cost, but a composite true-model decision-quality metric shows that repair is largely cosmetic on the frontier and dissociates from decision quality on the weak model, where a different cheap intervention delivers the largest gain instead. Separating estimation from execution error is what makes it visible that a plausible, cheap scaffold can degrade the axis it targets while leaving decisions unchanged.