Following Without Understanding: Target-Blind Attention Steering Before and After Instruction Tuning
Abstract
Post-hoc attention steering alters model generation by reweighting attention toward designated prompt spans, but existing methods typically select intervention components using downstream target-task evidence. We investigate what remains of instruction following when attention heads are profiled entirely without target-task prompts, labels, or reference responses. Using a target-blind narrative continuation objective on TinyStories (transferred), we find that the induced head ranking is remarkably conserved across pretraining and instruction tuning, whereas target-informed ranking (in-task) diverges across tuning. When applied to safety instructions on XSTest, the transferred heads induce strong, symmetric response polarization across both safe and unsafe requests. These results reveal a mechanistic dissociation: pretraining establishes a robust attention routing backbone that post-training preserves, but literal directive following operates independently of contextual understanding.