High Scores, Narrow Vocabularies: Measuring Cross-Trial Diversity in Large Language Models
Abstract
Large language models are increasingly used in creative industries as sources of novel ideas, treating output variation as evidence of genuine creative exploration. The Divergent Association Task (DAT), a validated creativity benchmark, has repeatedly produced at- or above-human level scores for LLMs, but a single-trial score cannot distinguish broad semantic exploration from repeated draws on a narrow, high-scoring vocabulary. We demonstrate this dissociation directly: LLMs can attain human-level or above-human DAT scores while showing substantially narrower cross-trial vocabularies than humans. We test five models, three open-source and two frontier, with 100 trials per condition under varied prompts, reasoning instructions, and decoding parameters, assessed using five metrics: DAT score, word overlap, vocabulary breadth, entropy, and semantic-region overlap. Prompt wording and reasoning instructions are the strongest predictors of DAT score, but leave lexical diversity largely unchanged. By contrast, widening the sampling pool reliably increases diversity without a corresponding gain in score for most models. Neither lever therefore closes the gap between benchmark performance and genuine exploration. Frontier models show the most extreme version of convergence, showing the highest cross-trial overlap and narrowest vocabulary of any model tested. These findings show that human-like DAT performance does not imply human-like diversity. We propose cross-trial diversity across multiple metrics as a behavioral diagnostic of generative convergence that single-trial scores cannot reveal. As LLMs become routine in creative work, such convergence may homogenize otherwise independent outputs, warranting caution in treating high DAT scores as evidence of genuine creative diversity.