VIPBench: A Human-Aligned Benchmark for Voice Identity Perception in the Age of Voice Cloning
Abstract
Speaker identification benchmarks evaluate how accurately a model can predict a ground-truth speaker label for a given utterance, treating voice identity as a property of the recording's producer. However, these benchmarks do not test whether a model aligns with human perception of voice identity. Whether a human would hear two voices and interpret them as the same speaker or not has consequences for intellectual property rights, cybersecurity, and everyday life. We make two contributions. First, we introduce the Voice Identity Perception benchmark (VIPBench), built from the identity judgments of English-speaking crowdworkers on stratified celebrity voice pairs. VIPBench consists of 124,876 same/different identity judgments from 1,290 participants on 9,800 voice pairs across 100 speakers, spanning real recordings, AI voice clones generated by a state-of-the-art Text-to-Speech (TTS) system, and continuously morphed voices. Our second contribution is to define four evaluation tasks: predicting the listener same-speaker agreement rate; binary same/different classification against the majority listener vote, evaluated for both ranking and calibration; alignment between the model's and listeners' speaker-similarity structures; and whether a predictor fit on real speech still works on voice clones and morphs. We report baselines for ten publicly available speech representations. We find that the perception target re-orders model rankings relative to metadata: supervised embeddings (trained on metadata speaker labels) still outperform self-supervised models (which learn voice structure without identity supervision), and no model reaches the noise ceiling. VIPBench enables systematic evaluation of speaker representations against human voice-identity perception across real, cloned, and morphed speech, motivating speaker models trained on perceptual judgments directly.