The Multilingual Artificial Hivemind: Do All LLMs Think Alike in Every Language?
Abstract
Text generated by Large Language Models (LLMs) often remains limited in diversity compared to that produced by humans. Recent advances demonstrate the presence of inter- and intra-model collapse in open-ended generation of English text, raising concerns about the long-term homogenization of LLM, and more importantly, human thought. However, our understanding of whether this homogenization extends to languages beyond English is limited. In this work, we investigate the role of language on output diversity. Specifically, we aim to understand whether multilingual LLM-generated text is also homogeneous, and whether this homogeneity matches what we observe in humans. Across 23 models and 13 languages, we find that the artificial hivemind effect is pervasive but uneven in its extent. The effect is not an artifact of how diversity is measured: it holds in three different embedding spaces and under a lexical metric that uses no embedding model. The size of the homogenization gap is predicted by a model's competence in a language, robustly across embedding spaces, but not by the language's tokenization fertility or morphological complexity. Output diversity is thus lowest where models are strongest.