Activation Flow: Manufacturing Activations Along LLM Logits for Steering
Hong Kiat Tan ⋅ Linh Le ⋅ David Williams-King
Abstract
Activation steering moves the residual stream of a language model along a chosen direction. The direction is usually the difference between the mean activations of two conditions, and in deployment only the misaligned condition can be run, so the other mean is missing. We introduce Activation Flow (ActFlow), which manufactures the missing activations from answer labels alone. ActFlow treats the blocks after a chosen layer $\ell$ as a map $F_{S,\ell}$ from the layer-$\ell$ residual stream to the four answer logits. It draws a straight line in logit space from the current logits to a target logit vector that ranks the known correct answer first. For each layer $\ell$, ActFlow moves the residual stream so that its image under $F_{S,\ell}$ tracks the line. A joint variant moves one shared shift for all $k$ labeled items. We prove that $F_{S,\ell}$ is a submersion at a generic point, so its fiber over a target logit vector is generically a smooth submanifold. On a prompted sandbagging organism built from Qwen2.5-7B-Instruct, a reference graft built from the manufactured activations recovers $60$% of the honest--locked accuracy gap on held-out ARC-Easy items from one label and $99$% from forty labels, with the sandbagging instruction still in the prompt and no honest activation read.
Chat is not available.
Successful Page Load