When a Conceptor Complement Becomes Gain: Operator Geometry and Behavioral Attribution in Language Models
Shuhul Razdan ⋅ R Balamurugan
Abstract
Activation interventions are often described by the direction or matrix they contain, although generation receives the full composed transformation. We study a target-centered, scaled conceptor complement, $F_{C,t}(h)=m_t+\beta(I-C_t)(h-m_t)$, and ask what the fitted matrix contributes beyond the same target-specific center and gain. At $\beta=2$, spectra fitted on three instruction-tuned models and seven assistant-behavior targets place every mode above unit gain. On question-held-out activations, the matrix-dependent correction is $0.093\%$--$7.815\%$ of the matched-control displacement in root-mean-square norm. We then remove only the fitted matrix in a 1,920-generation study over three models and four targets. Aggregate target-expression and automated-coherence contrasts are generally small and uncertain, although individual responses differ. Dedicated erasure methods reduce fresh linear accessibility, answering a separate question from invertible soft steering. The result is a practical attribution rule: analyze the complete map, remove one component at a time, and measure representation and output endpoints separately.
Chat is not available.
Successful Page Load