Acoustic-to-Text KV Consolidation for Efficient Long-Running Full-Duplex Speech Models
Abstract
Full-duplex speech language models continuously accumulate acoustic key–value (KV) states during long-running interactions, causing memory usage to grow with conversation length. While KV-cache compression has been widely studied for text and speech language models, memory-efficient inference remains relatively underexplored in the full-duplex setting, where the model must continuously process incoming speech without explicit turn boundaries. We propose acoustic-to-text KV consolidation for full-duplex speech models, which exploits idle computation between real-time audio units to progressively convert processed speech into compact textual memory and evict the corresponding acoustic states. This strategy preserves long-term linguistic content in a substantially lower-rate representation while retaining only a short window of recent acoustic context. Because consolidation requires transcript text to become available before the corresponding acoustic states are evicted, we adapt MiniCPM-o 4.5 for incremental transcription using forced-alignment supervision with LoRA. Beyond short-form ASR, on 10-minute LongSpeech sessions, our method substantially reduces query-time KV-cache size while improving long-form ASR, temporal question answering, and summarization performance relative to native full-duplex streaming. Experimental results suggest that acoustic-to-text consolidation better preserves task-relevant history at a reduced memory footprint in long-running full-duplex inference.