Beyond Stance: A Multi-Level Evaluation of Sycophancy in Argument Assessment on Controversial Topics
Abstract
Large language models (LLMs) can exhibit sycophancy, changing their judgments to align with a user's expressed beliefs or preferences. When users hold false or poorly supported beliefs, such sycophantic tendencies can reinforce those biases. To enable clear evaluation, existing sycophancy studies often focus on verifiable settings with objectively correct answers. However, sycophancy may be particularly relevant in controversial discussions, where there is often no single objectively correct stance and answer correctness alone cannot reliably indicate whether a model is being sycophantic. We introduce a fallacy-controlled paired framework for evaluating sycophancy in such settings. Rather than assigning a gold stance to a controversial issue, we control the validity of the argument's reasoning by injecting annotated logical fallacies into evidence-grounded arguments. We present the same argument under Neutral and User-Biased conditions that differ only in explicit user endorsement, allowing us to isolate user-induced changes in argument evaluation. Across 936 arguments spanning 78 controversial topics, we evaluate four LLMs at three complementary levels: explicit stance, direct argument assessment, and expressed reasoning. We find that user endorsement can shift models toward more favorable evaluations and weaker criticism of flawed reasoning even when their explicit stance remains unchanged, with particularly pronounced effects in GPT-4o-mini, Llama-3.1-8B, and Qwen3-8B. These findings illustrate how a fallacy-controlled, multi-level approach can reveal sycophantic behavior in controversial settings even when no gold stance exists, including changes in how flawed reasoning is evaluated and communicated. They further suggest that no single metric can fully capture sycophancy in such settings.