Selection on the Simplex: Language models implement persona selection through Bayesian belief geometry
Abstract
Language models have been frequently described informally as implementing ‘persona selection.’ Yet despite its wide usage for understanding language model behavior, this persona selection hypothesis has not been rigorously or directly tested—risking a system- atic misunderstanding, and missing out on the potential explanatory power that a formal model would provide. In this work, we define and test a formal model of persona selection as Bayesian inference over the weights of a mixture model, and evaluate the predictions of that formal model using behavioral and internals-based techniques. We find that (1) language models naturally learn multi-trait persona structure when finetuned on data gen- erated from those personas, (2) that their in-context belief updates about personas are approximately Bayesian, and become more Bayesian as we scale model and finetune corpus size, and (3) that they encode belief states linearly on a simplex in the residual stream, en- abling direct internals-based prediction and control over personas. Our findings constitute direct experimental evidence that language models implement a formal version of persona selection over multi-trait personas, connecting questions about language model personas to existing literature on Bayesian in-context learning, and taking steps towards a more general mathematical theory of language model behavior.