Stability Regimes for Framing-Sensitive Fine-Tuning in Language Models
Abstract
AI systems increasingly shape how individuals access and evaluate information, intensifying concerns about misinformation. Computational simulation offers a controlled complement to direct empirical study, with fine-tuned LLMs enabling powerful agent-based simulations. However, fine-tuning often induces broad behavioral shifts including bias drift, degenerate heuristics, and loss of task competence, which undermine experimental validity. The central challenge is inducing localized decision shifts without global behavioral collapse. We study this through a stability analysis of supervised fine-tuning, attenuating loss on selected training slices to shift false-claim acceptance while preserving overall task competence. We introduce FrameRef, a large-scale dataset of semantically equivalent claims with controlled surface framings across five dimensions (Authoritative, Consensus, Emotional, Sensationalist, Prestige), with human validation confirming systematic framing effects on claim acceptance. We identify stability regimes in which targeted error shifts can be induced without degrading aggregate accuracy or calibration. A sequential exposure task over FrameRef shows that small framing-conditioned shifts compound into substantially different cumulative outcomes under feedback, but largely disappear when feedback is removed, confirming that interventions modify specific decision tendencies rather than global performance. We release code, LoRA adapters, and data at https://github.com/anonymousauthor2352/frameref.