Acting Aligned: Topological and Geometric Analysis of Alignment-Faking and Emotion-Acting Representations
Abstract
We investigate shared representational structure between alignment-faking and acted versus naturally elicited emotion in LLMs using geometric and topological approaches. We use residual stream activations from Qwen3-32B, combined with persistent homology tools from topological data analysis (TDA) and simple linear analyses, to characterize differences in representations between alignment-faking under differing monitoring regimes and "emotion acting," finding distinct topological and linear separation in both tasks. We show cross-task generalization of probes in these settings, use rank-10 subspace analysis to show linear overlap above chance, and find that steering with a vector derived from emotion acting significantly influences alignment-faking behavior. Evidence from a "banal acting" control suggests that this overlap reflects representational similarity of acting behavior in general, extending existing literature results on role-playing and context-dependent representations.