Sandboxed Coding Agents are Competitive Omni-modal Task Solvers
Abstract
As multimodal LLMs increasingly emphasize video and audio, a common assumption is that solving such tasks calls for native omnimodal models. We show that this is not always necessary: coding agents equipped with only text+vision and a sandboxed tool-using interface can perform competitively with, and in several settings outperform, SOTA native omnimodal models and predefined multimodal agent scaffolds across multiple audio-video benchmarks. Our trajectory analysis suggests that their advantage comes from coding agents writing code and orchestrating tools to retrieve relevant content from transcripts, frames, and other non-textual modality signals from raw inputs. This effectively converts omnimodal tasks into evidence retrieval and information processing problems, avoiding the inefficiency of ingesting entire videos or audio streams into context. To characterize remaining limitations, we propose a failure taxonomy and a process-level analysis of tool-use traces, and find that simple skill injection, including human-written skills and self-distilled skills from execution logs, can substantially improve performance over the no-skill baseline. To examine whether such capability can be elicited in open-source models, we further introduce Code-X, a complete training recipe with the OmniCoding trajectory dataset and verifiable reward, providing an exploratory baseline on Qwen-3.5-9B and Qwen-3.6-27B. Finally, given the maturity of many-modality understanding, we argue that the more meaningful frontier lies in many-modality processing, and introduce TerminalBench-O, the first process-level benchmark designed for coding agents on real-world omnimodal processing tasks. Together, our findings open new directions for omnimodal content processing and evaluation.