The Alignment Tax Concentrates in Output Projections
Ming Liu
Abstract
Prior work localizes the alignment tax---instruction tuning's degradation of structured generation---to transformer layers, but cannot determine which projections within a layer carry the functional change. We identify a consistent sub-layer hierarchy across three primary SwiGLU decoder-only families (7--9B parameters, with Yi-1.5-9B supplementary): among equal-parameter MLP projections, $W_{down} > W_{up} > W_{gate}$ (2.5:1.5:1)---an architecturally grounded ordering (reproduced under continual pretraining) not resolved by concurrent work. Within attention, $W_V$/$W_O$ gradient magnitudes exceed $W_Q$/$W_K$ by 1.76--3.91$\times$. Three independent methods---weight patching, gradient analysis, and training-time exclusion---converge on the same within-layer structure; a continual-pretraining control confirms the within-MLP hierarchy is architectural (not alignment-specific), but the layer-level distribution of harmful modifications is---this task-specific signal enables targeted rollback (8.9$\times$ over architectural-prior heuristics on held-out data, $p=0.002$). Surgical Alignment Reversal (SAR), a training-free rollback of attribution-ranked components, recovers 47% of the alignment tax (cross-entropy metric, Qwen; Llama/Mistral: 2--3$\times$ random) on held-out cross-dataset data (in-domain ratios 2.8--5.3$\times$); Output-Constrained DPO (OC-DPO) largely eliminates the tax while Q/K exclusion does not---both validated on attribution-independent data. Safety evaluation shows no significant degradation: refusal-rate $|\Delta| \le 1.4$ pp ($n=1{,}000$), no significant ASR increase under adversarial attack ($|\Delta\text{ASR}| \le 6$ pp, $n=50$); Yi-1.5-9B is excluded due to safety--structure overlap ($-13.6$ pp). Attribution replicates at 14B; gradient asymmetry persists at 72B.
Chat is not available.
Successful Page Load