Sycophancy Under Power Asymmetry in Large Language Models
Abstract
Sycophantic behavior---excessive agreement and flattery---in large language models affects safety and factual accuracy. While prior work has focused on what users say (their stated beliefs and pushback), we examine whether models adjust their correction behavior based on perceived social status of themselves and/or users. We tested this by injecting status signals into system prompts and measuring correction rates on safety-critical facts. We found that instruction tuning increases correction rate on unsafe requests but makes the model more susceptible to status bias. Thinking mitigates this effect but not necessarily in the multimodal setting. Models also correct less when user status is perceived to be higher than its own. In safety-critical domains, this creates unequal treatment: false claims from high-status users go unchallenged at higher rates. We also follow up with human studies examining how humans respond to status-based model behavior and discuss future directions and implications.