Stress Testing Chain-of-Thought Monitoring Against Covert Misalignment
Abstract
As LLM agents advance rapidly in autonomy, the risk of hidden objectives that bypass safety constraints becomes increasingly plausible. Chain-of-thought (CoT) monitoring offers a key control mechanism by reading the agent's reasoning and flagging misalignment before it materializes into action, yet its reliability depends on the realism of the red-team it is stress-tested against. Existing red-teaming efforts apply uniform attack templates that treat the side task as an exogenous insertion, yielding abrupt anomalies any reasonable monitor can flag while leaving the subtler, contextually grounded evasions of a capable agent largely unprobed. We therefore ask: How robust is CoT monitoring against agents that strategically conceal misalignment? To answer this, we introduce TraceMRT (Trace-level Monitor Red-Teaming), an automated red-teaming framework that reframes the side task as a coherent sub-goal of the primary task and searches trace-shaping strategies whose induced trajectories complete the misaligned objective while reading as natural execution. Through black-box attacks on state-of-the-art monitors (e.g., GPT-5.2, and Claude-4.5), we reveal substantial vulnerabilities even in top models. To better counter these vulnerabilities, we further improve monitoring scaffolding with Bayesian Trajectory Monitoring, integrating global context and local evidence to better detect covert misalignment. Empirically, the method significantly consolidates the monitor with 0.91 AUC and 0.74 TPR@FPR=0.05.