Drop the Act: Probe-Filtered RL for Faithful Chain-of-Thought Reasoning
Swapnil Parekh ⋅ Naman Goyal
Abstract
Reasoning models often continue after they have already committed to an answer, producing reasoning theater: deliberative-looking steps that contribute nothing to correctness. We introduce ProFIL (Probe-Filtered Reinforcement Learning), a drop-in extension to Group Relative Policy Optimization (GRPO). Prefix counterfactuals identify post-commitment steps; a lightweight probe is trained once on activations of a base model and then frozen. During RL, the current policy produces rollouts, while a separate frozen copy of the base model teacher-forces the same rollout text to supply the probe activations. High-theater rollouts receive zero reward and zero policy advantage. Across GSM8K, LiveCodeBench, ToolUse, and MMLU-Redux and two model architectures, ProFIL reduces post-commitment theater by 11-100 %, raises faithful fraction (including $+24$ percentage points on LiveCodeBench under an independent GPT-4.1 judge), and shortens chains by 4-19% in three of four domains, while preserving or improving the reported task-accuracy metric. A matched length-penalty baseline worsens theater, isolating commitment detection from generic compression. Frozen-probe audits, an independent pairwise judge, and inference-time steering controls further support the result and address the RL-obfuscation concern. Probe weights, training configurations, evaluation code, and rollout caches are released for all four domains.
Chat is not available.
Successful Page Load