What a Fully On-Device Speech Agent Costs a Standalone XR Headset
Abstract
Benchmarking on-device conversational AI on idle hardware fails to predict real-world performance in extended reality (XR). Evaluating a complete speech-to-speech agent (Whisper -> Qwen2.5-1.5B -> VITS) on a Meta Quest 3S during active rendering reveals that concurrent inference severely degrades frame delivery regularity—increasing long frames from 1.09% to 26.94%—while average frame rates remain deceptively stable at 72.00 FPS. Furthermore, sustained language model decoding disrupts frame timing significantly more than bursty prompt prefill despite equivalent CPU duty cycles. We show that conventional mitigations like OS priority tuning and prompt chunking fail, whereas confining inference to the application's assigned three-core set reduces long frames by 1.22–1.68x depending on the baseline. However, this core placement strategy reveals a strict oversubscription cliff: exceeding three software threads collapses decode throughput by 27x. These findings establish that XR speech agents must be co-evaluated with rendering workloads, using frame regularity rather than average throughput as the primary stability metric.