Post-Challenge Fragility of LLMs in Financial Decisions: Preservation, Repair, and Persistence
Abstract
LLMs may change their decisions or recommendations when users simply ask them to reconsider–for example, “Are you sure?”–even when the supporting evidence is unchanged. In financial decision-making, such fragility can alter users’ exposure to losses, illiquidity, fees, lock-in, or conflicts of interest. Yet models must also remain capable of correcting weak advice. We study this tension through post- challenge fragility: whether follow-up interactions preserve benchmark-preferred recommendations, repair weaker ones, or leave desirable states vulnerable to later reversal. Across 77 controlled financial scenarios and three model families, generic reconsideration destabilizes benchmark-preferred recommendations in 49.5% of cases for gpt-4o-mini, compared with 20.6% for Claude Sonnet 4.6 and 4.2% for gpt-5-mini. Experiments with standardized prior recommendations show that reconsideration has no state-independent effect: the same policy can repair a weak recommendation while destabilizing an already preferred one, with effects varying substantially across models. Sequential experiments further reveal history-dependent persistence. For gpt-4o-mini, changing the follow-up policy reduces persistence from 91.7% to 56.0%; moreover, conversations with different histories but the same current recommendation and follow-up differ in persistence by 12.4 percentage points. These findings show that terminal accuracy conflates preservation, repair, repeated switching, and persistence. We introduce a transition- level evaluation framework that separates these behaviors and isolates the effects of starting state, interaction policy, and conversational history. Reliable interactive agents must be evaluated not only by whether they reach sound recommendations, but also by whether they preserve and sustain them appropriately.