Necessary but Not Sufficient: Auditing Same-Answer Counterfactuals for Chain-of-Thought Faithfulness
Abstract
Chain-of-thought (CoT) monitoring assumes that a model's written reasoning is informative about the factors that actually influenced its answer. We test this with an answer-conditioned counterfactual design: two prompts suggest different wrong answers, and among traces where the model adopts the first suggestion, we label whether the matched control would have changed it. Prompt-end linear probes predict this dependence label at AUROC 0.61-0.65 across three 7-9B models, beating every LLM text monitor we tried (0.51-0.62). But a battery of audits shows this is not evidence of hidden causal provenance. A baseline that reads no hidden state and no reasoning, only candidate plausibility and task metadata, reaches 0.76-0.84, and adding hidden states to it yields no lift; probes do just as well on prompts that never contained the hint; choosing hint candidates the model finds equally plausible collapses the signal; and interventions fail. Under stronger prompt injection the ordering reverses and text monitors win. Same-answer counterfactuals are necessary for evaluating causal disclosure, but not sufficient; we distill the audit into a checklist for probe-versus-monitor claims.