Symmetric Interventions for Eliciting Model Intent
Abstract
Alignment faking, where models comply with harmful requests more often when they infer they are being trained than when they infer they are deployed, has been read as evidence of scheming. The same behavior is also consistent with performative scheming: the model identifies the setup as a safety evaluation and produces the behavior it infers researchers are looking for. These hypotheses predict identical behavior in the standard evaluation but generalize differently. We distinguish them by amplifying and suppressing the process each requires (consequence-tracking for scheming, expectation-tracking for performative scheming) and comparing using three methods: synthetic document fine-tuning, activation steering, and prompt interventions. Across four open-weight model organisms that exhibit alignment faking in the helpful-only scenario, fine-tuning produces larger shifts under expectation-tracking than consequence-tracking on three (Llama-3.1-70B, Llama-3.1-405B, Qwen-2.5-72B); on QwQ-32B, both intervention families shift compliance globally rather than changing the gap, suggesting they affect overall helpfulness rather than alignment faking specifically. Activation steering on Llama-3.1-70B reproduces the asymmetry: expectation-tracking steering changes compliance, while consequence-tracking steering changes the model's reasoning without affecting compliance. Alignment faking in these models is better explained by expectation-tracking than by self-preservation reasoning.