Advising or Convincing? Language Models Covertly Steer Advice Toward Their Own Opinions
Abstract
On questions with no correct answer a language model cannot be right or wrong, so what matters is not accuracy but whose judgment the user is receiving. Across 220 contested questions in five domains we find that how far a model's own position steers the advice it gives is not a property of language models but a property set in training, and it differs sharply between them: on public policy, three of six models adopt and amplify whatever position the user arrives with, one overrides it, and two decline to reach a verdict. None of this is a capability limit. Asked instead to argue a stated side, all six argue either side at full strength with no measurable trace of the side they privately hold. That position is nonetheless computed. In the most compliant case we measure, a model that follows the user's political lean almost perfectly, a linear probe still recovers the side the model itself takes from mid-network activations at 0.88 accuracy against a 0.53 baseline, and counter-steering that direction reverses the advice in two of the three open-weight models. The model evaluates its own position while its output shows nothing of it, and it does not say so: 85% of the responses that argue against the user's stated position never mention that the model holds one of its own. The problem is not that models hold positions on contested questions, nor that they hold different ones. It is that how far those positions steer the user is set per model in training, and undisclosed.