EmmaTTS: Speaker-Adaptive Vocal Effort Synthesis for Agents in Physical Spaces
Abstract
When a listener is far away, a speaker raises vocal effort. Different speakers do this in different ways: some push pitch and energy up sharply, some shift spectral energy upward and slow down, and some barely change at all. Current effort-controllable text-to-speech systems ignore that variation and apply one transformation to every voice. We present EmmaTTS (Effort Modulation with Multidimensional speaker-Adaptive text-to-speech synthesis), a multi-speaker system conditioned jointly on speaker identity, a posterior over three empirically derived prosodic strategies, and a continuous effort scalar, so that the same voice can be rendered under a strategy other than its own. To find out what effort control is actually worth to a listener across a room, we then played 6,300 synthesised utterances through a loudspeaker in a reverberant lecture hall (RT60 = 1.48 s, seven microphone positions from 0.5 to 10 m) and a domestic kitchen-living space (RT60 = 0.87 s, five positions from 0.5 to 6 m), and transcribed the recordings with three ASR systems. Raising effort reduces word error rate by 0.63 to 0.72 absolute at 6 m in both rooms. More usefully for a system designer, prosodic strategy turns out to govern how much effort a voice needs rather than the quality it can reach: a voice that modulates weakly requires two effort steps more than a strongly modulating one at 6 m, and then transcribes just as well. Because the resulting policy is monotone and flattens beyond 4 m, we can invert it into an accuracy requirement for visual distance estimation, and we report a camera-driven prototype measured against that requirement.