Predictive but Not Steerable Until Scale: A Three-Axis Dissociation of Affect and Chain-of-Thought Faithfulness
Abstract
Prior work has shown that language models represent affective states like valence and arousal as linear directions in their residual stream, and that chain-of-thought explanations are often unfaithful to the model's actual computation. We test whether these two observations are connected. Using Qwen2.5-1.5B, 3B, and 7B on GSM8K, and adding Gemma-2-2B and Meta-Llama-3-8B as cross-family checks, we extract valence and arousal directions from residual-stream activations and test three claims. First, a logistic probe over these two directions predicts next-step faithfulness above chance at every Qwen2.5 scale (AUROC 0.726, 0.722, 0.590 at 1.5B, 3B, 7B; 95% percentile-bootstrap CIs excluding 0.5 at 1.5B and 3B) and generalizes to Gemma-2-2B and Meta-Llama-3-8B. Second, a Pearl-style screening-off test shows the affect signal carries information beyond surface difficulty, most strongly at 3B (+0.212 AUROC over a difficulty-only baseline). Third, activation-addition steering of the valence direction produces no reliable effect at 1.5B or 3B, but at 7B a low-dose intervention raises faithful-reasoning rate by 10.4 percentage points (95% CI [+1.9, +20.8]), with a bounded dose-response signature distinguishing it from a random-direction control. The predictive signal weakens at exactly the scale where causal steering emerges — a three-axis dissociation across scale, confound-control, and model family that we argue is a useful diagnostic for real-versus-spurious activation-steering claims.