Many Biases, One Circuit: Post-Training Tunes a Pretrained Cue-Routing Mechanism
Abstract
Language models often abandon a correct answer when a prompt cue points at a wrong one. This failure appears in sycophancy and several related cue-induced biases. We ask where this behavior comes from during post-training. Across six model and recipe lineages, these biases converge on a recurring reader-writer attention circuit that is already causally active before alignment. Generic assistant adaptation makes chat-formatted prompts engage this circuit, even without any sycophancy-targeted data. Later post-training stages then change how strongly the circuit influences the answer. Across public Tulu-3 and OLMo-2 checkpoints, the circuit's routing tracks each stage's change in cue-following, while causally raising the model's confidence does not reduce cue-following. Most importantly, swapping the naturally learned routing patterns of one training stage into another moves the behavior toward the donor stage while preserving uncued answers. The same circuit also contributes causally to two-turn retraction. Our results suggest that post-training does not build a new sycophancy mechanism. It changes when pretrained cue-routing machinery is used and how strongly it affects the answer.