What Should You Say? Linguistic Enrollment for Few-Shot Voice Personalization
Caren Han ⋅ Sangeun Lee ⋅ Ju Yeon Suk ⋅ Seungyeon Ji ⋅ Sang-Wook Yi ⋅ Hyunsuk Chung ⋅ Kyungreem Han
Abstract
Few-shot voice personalization can reproduce a target speaker from only a few seconds of reference audio, yet little attention has been paid to what the user should say. We formulate \emph{linguistic enrollment optimization} as a budget-constrained content selection problem. Using Korean as our primary testbed, we construct a linguist-reviewed 96-item lexicon annotated with phoneme, coda, morphophonological, and phoneme-transition information. A budget-aware greedy procedure selects complementary linguistic evidence per second without using downstream quality measurements. Under a matched five-second budget, high-utility enrollment increases speaker similarity by $0.06$ and reduces character error rate by $2.8$ percentage points. Human listeners also prefer high-utility enrollment across all four perceptual criteria, including $68\%$ versus $20\%$ for naturalness. Cross-lingual evaluation shows that enrollment sensitivity differs across Korean, English, and Mandarin, with the largest effect observed for Korean in the model evaluated here. Further analyses identify morphophonological and phoneme-transition coverage as particularly informative and show that content selection matters most under short recording budgets, while item ordering and component weighting have comparatively limited effects. Our findings demonstrate that few-shot personalization depends not only on how much reference speech is available, but also on what linguistic evidence it contains.
Chat is not available.
Successful Page Load