ECHO: Diagnosing Spatial Cue Access in Audio-Language Models
Daniel Huang ⋅ Alvin Shen ⋅ Kyle Xu ⋅ Aarush Rachakonda ⋅ Aryan Shrivastava
Abstract
Physical-world AI must turn raw sensor measurements into representations that support reliable reasoning. Stereo localization provides a compact test: left and right audio channels carry measurable evidence about source position. Traditional end-to-end accuracy reports only whether the final direction is correct, hiding whether a model failed to obtain the spatial cue or failed to use it. We introduce ECHO (Explicit Cue versus Heard Observation), a paired diagnostic framework, and StereoMusicQA, a benchmark built from real instrument recordings rendered at known positions. Together, they test whether a model can use a spatial cue that it fails to obtain through audio. Each of 1,045 test items has two views: cue-text provides the measured left-right level difference and conversion equation; audio-only provides labeled channels and requires the model to obtain a useful relationship from audio. Because cue-text also supplies the rule, this comparison measures the practical advantage of a written cue and rule, not the effect of one internal layer. For Qwen2-Audio, Gemini Flash-Lite, and GPT-audio, cue-text left/center/right accuracy is 71.9%, 93.6%, and 90.4% compared with 28.5%, 36.6%, and 42.0% audio-only. These 43.3–57.0 percentage-point gains remain positive when complete source tracks are resampled. Furthermore, a model can identify the correct side yet still predict an inaccurate angle: cue-text within-$5^\circ$ accuracy ranges from 0.4% to 80.7%, and written-calculation faithfulness ranges from 0.4% to 77.8%. ECHO shows that successfully using an explicitly supplied cue and rule does not guarantee that a model can obtain the required cue from sensor input, and that spatial evaluation should test both paths.
Chat is not available.
Successful Page Load