When LLMs Manipulate: Measuring Behavioral, Social, and Temporal Effects
Sadia Asif ⋅ Israel A Laguan ⋅ Shumaila Asif ⋅ Haris Khan ⋅ Muneeb Asif ⋅ Muhammad Zeeshan Karamat
Abstract
Large language models (LLMs) can influence user trust and decision-making through subtle linguistic cues, yet the behavioral effects of manipulative model outputs remain poorly quantified. We introduce a behavioral computing framework for measuring \textit{behavioral vulnerability}, defined as user susceptibility to manipulative AI responses. Building on DarkPatterns-LLM, we generate controlled manipulative stimuli using Manipulation Gradient Descent (MGD) targeting biases such as authority, urgency, and social proof. We analyze participant responses using the Multi-Stakeholder Impact Assessment Network (MSIAN) to model cross-user influence and the Temporal Harm Propagation Model (THPM) to characterize persistence over time. In a 120-participant study, manipulative framings produced significant changes in trust calibration and compliance ($p < 0.01$), with influence also propagating across social contexts. These findings demonstrate that LLM safety evaluation should extend beyond model outputs to measurable behavioral consequences and provide a computational framework for quantifying individual, social, and temporal risks from manipulative AI interactions.
Chat is not available.
Successful Page Load