Output Equivalence Does Not Imply Causal Invariance in Chain-of-Thought Monitors
Joe Smith
Abstract
Chain-of-thought monitors are expected to respond to what a reasoning trace says, not how it is phrased. We test that expectation at the output, input, and activation levels of one fixed monitor. Llama-3.1-8B-Instruct reads partial Qwen3-14B reasoning traces from 97 MMLU-Redux items, together with two independently checked meaning-preserving rewrites. The monitor's mean trajectory is equivalent across a light rewrite and a larger register shift: the area-under-curve difference is $-0.001,{\scriptstyle\begin{matrix}+0.011\-0.010\end{matrix}}$, within a prespecified $\pm 0.05$ equivalence margin. Yet replacing one aligned step changes policy-answer log-odds by $0.552,{\scriptstyle\begin{matrix}+0.062\-0.056\end{matrix}}$. Exchanging the final-token residual state between the two rewrites changes log-odds by $0.132,{\scriptstyle\begin{matrix}+0.015\-0.012\end{matrix}}$ at layer 8 and $0.336,{\scriptstyle\begin{matrix}+0.042\-0.033\end{matrix}}$ at layer 16. Full states from another item are more disruptive, even after matching answer, verdict, and confidence. That comparison does not identify a rewrite-specific direction: once difference vectors are matched in norm, directions from the same item, a matched donor, and an unrelated pair have similar effects. Scaling any of these vectors produces a monotone response while changing the monitor's argmax on at most 1.5% of trials. Stable verdicts can therefore conceal substantial internal sensitivity, and a patching effect alone does not show that the model uses the semantic distinction assigned to the patched direction. We release the code, monitor prompts, frozen preregistration, and item-level intervention records.
Chat is not available.
Successful Page Load