Extracting Steering Vectors from the J Space
Abstract
In this work we explore the idea of deriving steering vectors from a published Jacobian lens rather than from data. To that end, we run the lens backwards: the unembedding rows of a few concept tokens are pulled back through the lens Jacobian to the activation that would have verbalised them, contrasted against counter-tokens, and averaged. We evaluate the resulting direction against a difference-in-means vector fitted on the model over 8 behaviours on Qwen3-1.7B and show that it matches or exceeds the fitted vector on behaviours that amount to a preference over word choice. Finally, we present results suggesting that the method fails where a behaviour requires replacing the lexical stream: the derived direction encodes the topic of its concept tokens rather than the behaviour, and a projection of the fitted direction onto the subspace the tokens can reach fails on the same behaviours