Uncovering Hidden Propensities in Language Models via Limited-Parameter Finetuning
Abstract
Large language models can have hidden propensities: behaviours that arise only in rare contexts. These hidden propensities pose a challenge to alignment audits, which seek to uncover problematic model behaviour but are limited in scope to scenarios anticipated by the auditor. To address this challenge, we explore limited-parameter finetuning (LPF) on examples of a target behaviour. While training 3–15\% of model parameters, LPF reveals whether the model has a hidden propensity, even when standard behavioural evaluations cannot. LPF succeeds on a suite of model organisms with hidden propensities (8B-14B parameters), and also surfaces naturally occurring hidden propensities in open-weight models (27-32B parameters). LPF can thus distinguish between behaviourally equivalent models, even without knowing the context in which the propensity surfaces. LPF provides a novel affordance that may be useful for auditing frontier AI models.