Within-Model vs Between-Prompt Variability in Large Language Models for Creative Tasks
Abstract
When a large language model (LLM) generates a creative response, how much of the outcome is determined by the prompt, the model, the task, or pure sampling luck? We address this question with a variance decomposition over 89,806 generations from 10 LLMs, 6 prompt strategies, and 30 Alternative Uses Task items, which is one of the most widely used creativity tests. For output quality (originality), the results invert common assumptions: the AUT item being evaluated and the prompt strategy together account for over 72% of variance, while model choice contributes less than 8% and within-LLM stochasticity alone exceeds model differences by more than 2×. Controlling for the AUT item sharpens the picture further: prompt strategy then explains 3.4× the variance of model choice. For output quantity (fluency), the pattern reverses, with the LLM being the dominant factor. We further show that prompts shape output distributions, not point estimates: “discriminative” prompts (e.g., Persona) widen the quality distribution and show higher peaks in capable models, while “constraining” prompts (e.g., Format) compress it. Together, these results argue that single-sample LLM evaluations are unreliable for open-ended tasks, and that evaluation designs must account for the full variance structure of generative systems.