Reconstructing the Vocal Tract with Differentiable Acoustic Simulation
Abstract
The human vocal tract, the cavity consisting of one's throat, mouth, lips opening, etc., filters one's voice to create the sounds we know as speech. In this paper, we present a differentiable and GPU parallelizable acoustic simulator that synthesizes speech by propagating sound along an acoustic tube representation of the vocal tract, and via its gradients, solves the inverse problem: reconstructing the geometry of their vocal tract solely from the sound it produces. Although learning the inverse mapping between geometry and sound is notoriously non-convex, we discover that gradient descent succeeds with three technical contributions: (1) we design a frequency domain formulation of the vocal tract's fluid dynamics that is 70x more GPU parallelizable than finite differences in time, (2) we integrate a differentiable model for turbulence to synthesize consonants, and (3) similar to prior work in implicit neural representations (INRs) and neural radiance fields (NeRFs), we find that parameterizing the geometry with a neural network accelerates convergence and escapes local minima that trap discrete representations. Because the simulator is differentiable, it is readily integrated with other deep learning pipelines to enable novel speech and medical imaging applications. Our model can be used for instance, in singing instruction or language learning, where a visualization of one's vocal tract can help people understand how their vocal tract maps to different speech sounds. To enable these tasks, (1) we demonstrate self-supervised autoencoding of vocal tract shapes across 11 languages, and (2) we couple our simulator with a generative model of MRI images to reconstruct one's moving vocal tract from only their speech, no paired data required.