Clarification Is Not Correction: LLMs Fail to Let Go
Abstract
Dialogue failures in language models are often framed as memory failures: the context is too long, the summary is lossy, or the model has forgotten an earlier constraint. We argue that this framing misses a deeper problem. In many conversations, the model does not simply forget; it commits too early. An ambiguous early turn is collapsed into a single hidden interpretation, and later clarification is filtered through that commitment. We call this failure mode as early posterior collapse: the collapse of unresolved user intent into a committed task state before ambiguity has been resolved. We study this phenomenon through controlled dialogue tasks in writing, planning, and coding, using Gemini-2.5-Pro and Gemini-2.5-Flash. Across thousands of trials, we find that the same information presented in different orders leads to different downstream outcomes, even when the final dialogue contains equivalent task-relevant information. This order effect suggests that later clarification is often treated as additional context rather than as a corrective signal: it refines a stale task state without necessarily invalidating it. Coding tasks are especially vulnerable, suggesting that early assumptions become embedded in structured artifacts such as interfaces, constraints, and control flow. Standard prompting and memory strategies do not reliably solve the problem: summaries can collapse ambiguity, and chain-of-thought can reduce explicit wrong commitment in reasoning traces without reliably improving final task success. These findings motivate uncertainty-preserving state management. If assistants fail to let go of early interpretations, robustness cannot rely on post hoc correction alone; it must also prevent ambiguous early turns from hardening into a single task state. Assistants should maintain tentative hypotheses while ambiguity remains, ask before executing when high-impact ambiguity persists, and rebuild from a revised task state when later evidence invalidates an earlier interpretation. Rather than proposing a single prompting fix, our goal is to redirect robustness research for interactive LLMs from retaining more context toward preserving uncertainty until clarification can operate as correction.