Probing Human–Machine Perceptual Alignment with Semantically Ambiguous Images
Abstract
The classic duck–rabbit illusion reveals a fundamental property of perception: when visual evidence is ambiguous, a visual system must resolve competing interpretations. But do humans and machine classifiers resolve the same ambiguity in the same way? To explore this, we use semantically ambiguous images as probes of human–machine perceptual alignment. Our psychophysically-informed framework interpolates between concepts in vision-language embedding space to generate continuous spectra of ambiguous images, enabling direct comparison of human and machine perceptual decision boundaries and sensitivities. We find systematic differences in these boundaries: machine classifiers are shifted relative to both human observers and the midpoint of the synthesis continuum. Guidance scale also produces larger changes in human perceptual sensitivity than in the machine classifiers tested. Controlled semantic ambiguity in generated images, combined with a psychophysical framework, can thus enable quantitative evaluation of human–machine perceptual alignment.