Mirroring Is Not Sycophancy: Style Convergence and Position Change in Human--AI Conversation
Abstract
Preference-based alignment relies on human judgements that are assumed to be independent of the model being judged. We test that assumption by measuring how each party in a human–model conversation changes the other. In a sample of 25,183 conversations from a 125,464-conversation evaluation corpus, users adopted the vocabulary of the model they were talking to at a rate above chance, with chance estimated from conversations that began with the same question. The amount of adoption predicted the user's eventual preference vote about as well as reply length did, and the prediction was already available from the user's first three turns. In a scripted counterfactual audit of 16 models, the user's expressed stance shifted every model's substantive position and the user's emotional register shifted every model's style. These two effects were unrelated across models. The style ranking from the audits matched the style ranking from the real conversations (ρ = 0.68), but neither predicted which models changed position (ρ = −0.14). Position changes did not transfer to an unrelated pressured question two turns later. Replacing the human with a second model in 699 further conversations all but removes the asymmetry, and one sentence telling a model its partner is not a person removes most of its deference. Preference votes depend on the user's own convergence, though whether that is a bias in the vote or an early signal of quality, our data cannot separate. What the corpus and the audits agree on is clear: mirroring is not an indicator of sycophancy.