Speaker Activity for Generative Speech Separation: Conditioning, Selection, and the Limits of Gradient Steering
Francisco Sumba Toral ⋅ Haolong Zheng
Abstract
Conversational agents need speaker-attributed audio when speakers interrupt or talk over one another. We examine three uses of predicted speaker activity in a frozen flow separator: prompt conditioning, candidate selection, and gradient steering. On 2,640 paired mixtures, diarizer-derived temporal anchors improve frozen reference-free selection by $0.645$ dB and oracle best-of-16 quality by $0.663$ dB relative to a generic prompt, with gains concentrated below 80% overlap. In an eight-mixture mechanism audit, a frozen diarization reward identifies substantially better candidates, but five gradient updates increase reward while changing SI-SDR by only $-0.04$ dB. A 20-mixture DiffSep proxy with a normalized-energy reward shows the same ordering: reranking outperforms terminal optimization and one-step steering. Two small locked follow-ups find that a conditional energy update can rescue collapsed outputs, whereas a learned halfway critic does not significantly outperform a matched-compute energy selector. These results separate the benefits of conditioning and selection from the failure of the tested activity gradient. The study is audio-only and offline; $K=16$ sampling costs 3.52 aggregate GPU-seconds per audio second, and end-to-end latency was not measured.
Chat is not available.
Successful Page Load