Scalable Search Heuristics for Inoculation Prompting
Avyukth Nilajagi ⋅ Xinyue Liu ⋅ Nikhil Maturi ⋅ Bhiman K Baghel
Abstract
Inoculation prompting (IP) defends against undesired behavioral generalization by fine-tuning models with prompts that explicitly request the undesired behavior. However, selecting an effective IP requires a per-candidate fine-tuning run, making prompt search computationally expensive. Previous work proposes that the degree of undesired behavior induced by the prompt on the base model (behavioral elicitation rate; BER) is directly correlated to that prompt's effectiveness as an inoculator. We show this correlation does not always suffice to construct a meaningful ordering of candidate IPs with respect to effectiveness. We hypothesize that the model's latent space provides a more fine-grained distinction of an IP's effectiveness. To this end we introduce $\mathbf{La}$tent $\mathbf{P}$reference $\mathbf{S}$hift $\mathbf{(LaPS)}$, a heuristic that predicts an IP's effectiveness from base-model log-probabilities. LaPS computes the difference in log-probabilities for pairs of undesired \& desired behavior-expressing responses. LaPS consistently improves upon BER in the reward-hacking and sycophancy settings, particularly when behavioral elicitation saturates. Pearson $r$ is 0.667 vs. 0.473 for sycophancy and 0.940 vs. 0.287 for reward hacking, while Spearman $\rho$ is 0.639 vs. 0.487 and 0.867 vs. 0.389, respectively (LaPS vs. BER). Here, correlations are computed between heuristic predictions and the post-finetuning effectiveness of candidate IPs. LaPS achieves consistently higher point-estimate correlations than BER across all evaluated settings.
Chat is not available.
Successful Page Load