Modeling the Vividness of Imagined Natural Scenes Reveals a Model-Common Image-Level Component in Vision Models
Abstract
The visual features that support vivid mental images remain poorly understood. Here, we collected more than 229,000 vividness judgments for the 73,000 images in the Natural Scenes Dataset in a large online experiment (n = 1,991 participants). On each trial, participants viewed two images, were cued to mentally recreate one of them for 4 seconds and then rated the vividness of their mental image on a continuous scale from 0 to 100, followed by a memory task to encourage compliance. Using these data, we trained lightweight MLP readout heads on frozen vision-model representations to predict imagery vividness directly from image content. Several models predicted human vividness judgments above a low-level features baseline, with the best individual model reaching r = 0.376 [0.355, 0.396]. A single principal component of model predictions captured most of the cross-model variance and matched or exceeded individual models in predicting human judgments (r = 0.378). Even after removing human-aligned variance, residual predictions remained highly correlated across architectures, indicating a strong shared structure in model predictions. Attempts to extract additional human-aligned signal through residual probes, fresh layer-cached probes, and fine-tuning failed; fine-tuning instead further amplified the shared component. This shared axis also correlated more strongly with image memorability than with vividness itself, while observer-aware models generalized poorly to held-out raters, suggesting that models rely on a generic image-affordance signal rather than a vividness-specific representation. Together, these results show that subjective imagery vividness is partially predictable from visual content alone, while revealing a systematic gap between predictive performance and construct-specific representation in current vision models.