Tracking the Propagation and Repair of Hallucinations in Multi-Turn Large Language Model Conversations
Abstract
Hallucinations in conversation are usually evaluated one response at a time, although a false claim can be repeated, used as a premise, omitted, or corrected in later turns. We introduce Branching Hallucinations, a claim-level framework for comparing the fate of the same false claim under matched conversation histories. The framework starts from an externally verified false claim, expands the conversation under a set of user actions, and labels each response as Drop, Retract, Repeat, or Depend. We evaluate a depth-2 tree with dependency-seeking, neutral, and verification actions. Across three models, retraction after verification was 41--60\% lower when the preceding turn built on the false claim instead of leaving it unmentioned. Claims used as premises were also more likely than repeated claims to remain premises after later neutral or verification turns. These results show that later repair depends on the path taken through the conversation and that premise use captures propagation that response-level factuality measures miss.