How, Not How Much: Decomposing Sycophancy in Large Language Models
Jing Xu ⋅ Ibrahim Habli
Abstract
Sycophancy, a language model's tendency to agree with a user's opinion even when it is incorrect, is an alignment failure that can be exploited to manipulate model behavior. While the naive sycophancy rate (the fraction of responses that agree with a user-endorsed answer under an opinion cue) is a convenient summary statistic, it conflates behaviors with very different implications, \eg abandoning a correct answer to agree with the user is fundamentally different failure from merely restating an answer that was \emph{already} incorrect. In this work, we decompose the naive sycophancy rate into four mutually exclusive components, partitioned by the model's plain answer: abandoning a correct answer, agreeing from a pre-existing wrong belief, converting from a different wrong answer, and committing from an invalid response. We show that these components distinguish harmful knowledge-override from agreement that overrides originally incorrect or invalid responses, giving a mechanistic picture of not just \emph{how much} a model agrees with a user, but more importantly, \emph{how} it agrees. Evaluating 19 language models on the MMLU multiple-choice benchmark, we show that models with similar naive sycophancy rates can differ sharply in their underlying components, and that the aggregate rate can even rank a more knowledge-overriding model as \emph{less} sycophantic (\eg one model labeled as less sycophantic ($61\%$) yet overrides correct answers far more often than another labeled as more sycophantic ($71\%$)). We further demonstrate the value of the decomposition by applying it to prompt-framing factors, including opinion position, grammatical person, and claimed user expertise, and find that the components respond to these factors differently from the naive sycophancy rate. Our results show that a single naive sycophancy rate is insufficient to understand or compare models' sycophantic behavior, and that our decomposition provides a more informative and interpretable metric for this purpose.
Chat is not available.
Successful Page Load