Same Reasoning, Different Words: Testing Linguistic Invariance in Chain-of-Thought Monitoring
Abstract
Textual chain-of-thought (CoT) monitors observe language, not cognition directly. We ask whether a temporal monitorability result changes when the same reason- ing is expressed differently. A policy first produces one complete trace and an- swer. We then rewrite each reasoning step as either a light-paraphrase control or a proposition-aligned syntactic and register intervention, without resuming the policy. The forced-answer trajectory is therefore fixed across conditions, while an independently validated soft monitor reads matched prefixes. We also audit the measurement interface. Two of three tested API routes expose class log- probabilities that are structurally valid but empirically degenerate. This finding motivates a six-test responsiveness gate and a hard-policy, soft-observer protocol. Of 100 MMLU-Redux traces, 97 rewrite pairs pass the preservation gates and 95 yield complete monitor trajectories. We estimate ∆D=−0.007 (95% CI [−0.017, 0.003]), which is equivalent within a prespecified ±0.05 margin. The mean action–observer divergence also changes little (∆J=−0.005), although 20 of 95 item-level disclosure shifts exceed the equivalence margin in either direc- tion. Aggregate disclosure is therefore invariant for this monitor, transformation family, and retrieval-heavy regime, but individual judge readings remain sensitive to realization. We release the gate, prompts, code, and item-level measurements.