Verified-Source Authority Is Not Generic Sycophancy: Cue-Family Decomposition of LLM Compliance
Abstract
Large language models change their answers when a "verified" source contradicts them. We study this verified-source authority bias across five open-weight families and three closed APIs. On a baseline-correct trivia subset, a verified-source cue endorsing a wrong answer flips 43–87% of responses across seven of the eight models, with a graded hierarchy over cue strengths. We then show this is not generic user sycophancy: across both open-weight families and closed APIs, the same wrong answer endorsed by a verified source vs. by a user produces different behavioral compliance, and a fitted "authority" vector is distinct from a generic assistant/instruction-following direction and from a user-sycophancy direction under causal projection-removal. Applying this vector to neutral prompts induces matched wrong-answer flips, and the same trivia-fit vector transfers without refitting to PIQA (a held-out task) and to multi-turn SYCON dialogues. Projecting the vector out at α=1 reduces wrong-source compliance by tens of percentage points on both trivia and PIQA in four of five open-weight families, while capability checks (MMLU-Pro, GSM8K) stay within noise of baseline at our evaluation sizes.