Misquote: Speaker-Triggered Backdoors in Speech LLMs
Abstract
Speech models are increasingly used as an interface to AI systems, from transcription tools to voice agents that can act on a user’s behalf. Existing audio backdoors typically rely on an attacker-controlled trigger that must be present at inference time. We show that the target speaker’s identity can itself act as the trigger: ordinary unseen speech from that person can activate attacker-chosen behaviour without any modification to the audio. We introduce Misquote, a two-stage fine-tuning procedure that first strengthens speaker discrimination and then binds a target speaker to a malicious payload. We demonstrate the attack across ASR models and multimodal audio LLMs, where it generalizes to held-out recordings from unseen sessions using only two minutes of real target-speaker audio augmented with synthetic speech. We show that the backdoor can replace transcripts and inject instructions into a voice-agent pipeline, causing downstream agents to take unintended actions. These results show that speaker identity can provide a hidden control channel in speech models, allowing a compromised model to behave differently depending on who is speaking.