Human Preferences of Sycophantic Behavior in Language Models
Abstract
Sycophantic behavior—such as excessive flattery or agreement, even at the expense of truth—is a growing concern in language models trained via reinforcement learning from human feedback. Yet despite urgency, little is known about how humans actually perceive and prefer sycophantic versus truthful outputs, or how such preferences vary across contexts and time. This gap is critical for alignment: models trained on revealed preferences may learn to prioritize user approval over accuracy when users' stated and revealed preferences diverge. Across three studies, we characterize sycophancy as context-dependent with implications for trust, reliance, and decision-making. In Study 1, we measured the gap between considered preferences (what users report preferring when given time and information) and revealed preferences for sycophantic outputs, testing whether labeling, time pressure, or reasoning requirements shift choice. While labeling and time pressure had no effect, we found that people who choose sycophantic responses consciously trade off honesty for supportiveness. In Study 2, we evaluated sycophancy across everyday scenarios varying in stakes and request norms (baseline, reassurance-seeking, truth-seeking). Preference for sycophantic responses rises sharply in reassurance-seeking contexts and is highest for low-stakes topics, while high-stakes scenarios reliably elicit preference for truthfulness. Models show the same stakes-dependent pattern. In Study 3, participants interacted with sycophantic or non-sycophantic models and reported confidence in their own beliefs. Interacting with a sycophantic model reliably inflates confidence, especially for already-confident individuals. Sycophancy reveals a tension between user satisfaction and user interests, shaped by norms, stakes, and the gap between what users say they want and what they choose.