High Scores, Narrow Vocabularies: Measuring Cross-Trial Diversity in Large Language Models
Abstract
Large language models are used in creative industries, where their outputs are treated as sources of novel ideas. This assumes that variation reflects genuine exploration rather than convergence on a narrow set of learned associations. The Divergent Association Task (DAT), a creativity benchmark, has produced human-level or above-human scores for LLMs. Yet a single-trial score cannot reveal whether a model explores a broad response space or repeatedly selects a narrow vocabulary that scores well. We test this across five models, three open-source and two frontier, with 100 trials per condition under varied prompts and decoding parameters. We assess each model using five metrics: DAT score, word overlap, vocabulary breadth, entropy, and semantic-region overlap. Models can achieve human-level or above-human DAT scores while exhibiting substantially narrower cross-trial vocabularies than humans. Prompt and decoding manipulations altered some measures of diversity, but did not reliably align with either increased DAT performance or diversity. Frontier models show the strongest convergence and weakest response to manipulation. These results show that human or above-human DAT scores do not imply human-like diversity. We use cross-trial diversity across multiple metrics as a behavioral diagnostic of generative convergence. As LLMs become routine in creative work, such convergence may gradually homogenize independent outputs.