Agents Retrieve but Do Not Re-Derive: Coherence Costs Concentrate on Conclusion
Abstract
When multiple LLM agents share a memory store, one agent's action can silently invalidate another's cached entry. We ask: does this staleness change downstream behavior? Through 6,000+ counterfactual trajectories on three model families (GLM-5.3, Qwen-3.5-9B, Claude Opus 5), we identify a sharp structural boundary. If the stale entry states a premise (e.g., "tier changed to silver"), the agent ignores it and follows the stale conclusion 91 to 100% of the time on mid-scale models. If the entry states the conclusion itself (e.g., "current discount is 5%"), the agent corrects immediately, regardless of whether the correction is buried in prose or surfaced as a structured field. The effect is 98 percentage points, replicates across two task domains, and persists under efficiency pressure, metadata enrichment, and explicit verification instructions. We formalize this as a re-derivation gap: agents retrieve conclusions from memory but do not re-derive them from available premises. We prove that under this gap, standard coherence protocols (write-invalidate with evidence append) have zero expected benefit on conclusion-bearing entries, while a conclusion-rewriting protocol guarantees correction. We propose GuardWrite, a two-stage memory protocol that (1) attaches machine-checkable validity predicates to entries at write time, producing safe refusal when violated, and (2) rewrites the entry with the corrected conclusion when available, producing correct action. End-to-end, the protocol achieves 100% correct action after rewrite, compared to 100% stale-following without it. A frontier model (Claude Opus 5) closes the gap partially: it refuses to act on stale premises (0% stale-following) but still cannot re-derive the correct conclusion (0% fresh action on premise cells), confirming that the gap narrows from "wrong action" to "safe refusal" with capability but does not vanish.