Virtual-Flow: Virtual Microphone-based Speech Enhancement via Unsupervised Flow Matching
Abstract
Neural network-based virtual microphone estimation (neural-VME) synthesizes microphone signals at unobserved spatial locations from a limited set of real microphone (RM) physical recordings. This enables high-resolution arrays that improve spatial processing for downstream tasks such as speech enhancement. However, existing methods rely on supervised training with ground-truth signals captured at virtual locations, requiring datasets that scale with the number of virtual microphones (VM) and leaving open the question of how to design optimal large-array configurations. We propose \textit{Virtual-Flow}, a flow-matching framework that implicitly generates VM signals conditioned on real recordings, without requiring ground-truth targets. We introduce an unsupervised training strategy based on pseudo-VM targets constructed from time-delayed superpositions of real recordings, and show that it improves magnitude estimation in high-frequency bands, yielding richer spatial cues for beamforming. Experiments demonstrate that the proposed Virtual-Flow, combined with reflow distillation, outperforms existing state-of-the-art neural-VME models on downstream neural beamforming and speech enhancement tasks while requiring lower computational load.