Dual-State KV Forking: Low-Overhead Semantic Interruption for Full-Duplex LLM Agents
Akshat Sharda ⋅ Zhenyu Zhang
Abstract
Human conversation is full-duplex: listeners process speech continuously and interrupt mid-utterance to correct errors, volunteer answers, or take the floor. Multi-agent LLM systems remain half-duplex: a listener must wait for the speaker's entire turn, and input streaming—while it keeps the listener's cache warm—remains passive. We present the Dual-State KV-Forking architecture, which gives a streaming listener active floor control. The listener maintains a progressively updated Base KV cache over the incoming stream; at each chunk boundary it forks this cache (0.01–0.06 ms measured) into an ephemeral Assessor branch that appends a short decider prompt and emits a single logit-gated STOP/CONTINUE token, leaving the Base state intact. Each stream token is ingested exactly once, eliminating the $O(N^{2}\cdot k)$ context re-prefill of a stateless independent decider while preserving a warm cache for instant response generation. Across three models on physical hardware, on FLEXI, Full-Duplex-Bench, and a 200-scenario progressive-clue floor-control task, our framework matches or exceeds independent-decider accuracy (up to 100% on FDB), cuts total tokens processed per scenario by 78.6%, reduces Time-to-Halt by up to 49.1% under concurrent full-duplex streaming, and cuts no-interruption Time-to-First-Token by up to 65.0%.
Chat is not available.
Successful Page Load