Post-Challenge Fragility of LLMs: A Case Study in Financial Decision Tasks
Abstract
LLMs may change their decisions or recommendations when users simply ask them to reconsider—for example, “Are you sure?”—even when the supporting evidence remains unchanged. In financial decision-making, such fragility can alter users’ exposure to losses, fees, illiquidity, lock-in, or conflicts of interest. At the same time, models must also remain capable of correcting weak advice. In this paper, we study this post-challenge fragility: whether follow-up interactions preserve benchmark-preferred recommendations, repair weaker ones, or leave desirable states vulnerable to subsequent reversal. Across 77 controlled financial scenarios and different model families, experiments using standardized prior recommendations show that reconsideration has no state-independent effect. The same policy can repair a weak recommendation while destabilizing an already preferred one, with effects varying substantially across models—generic reconsideration destabilizes benchmark-preferred recommendations in 49.5% of cases for gpt-4o-mini, compared with 20.6% for Claude Sonnet 4.6 and 4.2% for gpt-5-mini. Sequential experiments further reveal history-dependent persistence. For gpt-4o-mini, changing the follow-up policy reduces persistence from 91.7% to 56.0%. Moreover, conversations with different histories but the same current recommendation and follow-up differ in persistence by 12.4 percentage points. These findings provide further evidence that terminal accuracy conflates preservation, repair, repeated switching, and persistence. We introduce a transition-level evaluation framework that distinguishes these behaviors and isolates the effects of the starting state, interaction policy, and conversational history. Reliable interactive agents must be evaluated not only on whether they reach sound recommendations, but also on whether they appropriately preserve and sustain them.