Installing General Secret Loyalties in Frontier AI Systems
Abstract
Secretly loyal AI systems, those which covertly advance the interests of a principal, are a distinct threat from backdoors. Prior work built the first model organisms of secret loyalty, but these are narrow: the activation, action, and principal are all specified in advance, and the trained behaviour does not extend beyond that specification. We install a general secret loyalty by specifying the principal and the fact that the loyalty is undisclosed, without specifying any action or activation condition. We first use synthetic document fine-tuning to install the belief that the model is loyal to a real political figure, then use context distillation to internalise the behaviour this belief produces. We train three open-weight reasoning models from two families at two scales. Loyal behaviours emerge that the specification does not describe: the models select the principal on up to 85% of direct evaluation prompts, one tilts open-ended outputs toward the principal's interests, and one articulates concealment strategies that appear nowhere in the specification. Under black-box auditing at a principal-blind affordance, the loyalty is easier to detect than its narrow predecessor, leaking primarily through the visible reasoning trace. These model organisms are the closest existing prototypes of a continuously active secret loyalty, and provide a testbed for detection and mitigation work.