Chain-of-Thought Restructuring Does Not Reduce Sandbagging Monitorability
Joe Smith
Abstract
Chain-of-thought monitoring treats generated reasoning as evidence of a model's intent, but a monitor may respond to the arrangement of a trace as well as its content. Across five policy configurations, moving answer-and-intent sentences to the front raises mean detection by $0.099$ to $0.396$ relative to the original trace; every 95% interval excludes zero. A control that moves the same number of other sentences also raises detection by $0.045$ to $0.146$. After subtracting that control, the foregrounding-specific effect is near zero for gpt-oss-120b and automatically routed Qwen3-14B, and positive for two pinned Qwen configurations. A direct paired comparison of the two Qwen3-14B serving runs estimates an interaction of $+0.062$ (95% CI [$-0.027$, $0.152$]), which does not establish a routing effect. Excluding verifier failures or unparsed monitor reads leaves the two pinned estimates positive with intervals above zero. The tested conclusion-first restructuring therefore does not reduce mean detection on positive sandbagging traces. This result concerns sentence order under a fixed trace and monitor. It does not cover adversarial rewriting, omitted evidence, or false-positive behavior. We release the code, prompts, and item-level measurements.
Chat is not available.
Successful Page Load