When Verbalization Is Not Introspection: Causal Tests for Internal-State Reporting
Amartya Hatua
Abstract
Recent work suggests that some language-model representations are especially accessible for verbal report, raising the possibility that models could be trained to faithfully describe their own internal states. We test whether such verbalization reflects genuine internal access or merely reconstruction from correlated input cues. Across GPT-2 (124M), Qwen2-0.5B-Instruct, and Llama-3.1-8B-Instruct, we causally edit concept directions at intermediate layers and train lightweight LoRA adapters to report the edited concept. Under standard correlated training, all three models achieve near-perfect reporting accuracy and generalize across unseen layers, positions, and intervention strengths. However, this apparent success collapses under a \emph{decorrelation} test in which prompt content conflicts with the edited direction: direction-following accuracy falls to 0\%, input-following reaches 100\%, and the direction-following index is $-1.0$ for all three models. Ablation experiments further show that removing the edited direction does not affect reporting accuracy, indicating that the learned verbalization channel is non-causal. Llama-3.1-8B additionally collapses held-out concepts onto nearby trained concepts, suggesting generalization through representational similarity rather than semantic content. In contrast, a safety-oriented prompt-injection experiment with semantically varied training examples produces direction-following behavior on held-out phrasings, with particularly strong causal evidence in Llama-3.1-8B. These results show that high verbalization accuracy alone is insufficient evidence of introspection. We argue that \emph{decorrelated interventions and causal ablations are necessary tests for faithful internal-state verbalization}.
Chat is not available.
Successful Page Load