AudioLoop: Program-Guided Iterative Reasoning Over Long-Form Speech
Abstract
Current audio-language models (ALMs) degrade on long-form speech even when recordings nominally fit their context windows. We introduce AUDIOLOOP, a task-conditioned, program-guided framework in which a root language model writes code to segment a recording, invoke localized audio or text sub-models, and accumulate evidence in a persistent execution environment. Across multi- hop debate question answering, contradiction detection, prosody trajectory estimation, and open-domain speech question answering, we compare memory persistence, segmentation strategy, and input modality. Persistent and context-aware configurations outperform stateless uniform processing, and direct audio reasoning is most valuable when the task depends on prosody rather than lexical content. Compared with single-pass full-context inference, structured iterative processing is competitive on multi-hop reasoning and substantially stronger on prosody trajectory estimation and contradiction detection, while exposing localization and evidence aggregation as controllable interfaces. We additionally release IQ2-QA, 39 multi-hop questions annotated over existing Intelligence Squared debate recordings. These results position external memory and audio tools as a practical interface for long-context speech reasoning; extending the framework to reusable memory and open-ended queries is the natural next step.