EchoProps: Can Multimodal Foundation Models Hear Physical Properties?
Abstract
Impact sounds carry the physical properties of their source: modal frequencies reflect size and geometry, and decay rates reflect material. Audio-capable foundation models are evaluated on semantic content or on signal-level attributes such as pitch, not on whether they can recover the properties of the object that produced a sound. We introduce EchoProps, a benchmark for object-level physical property inference from impact sounds. It comprises material identification and size ordering on real measured recordings, a single-parameter damping sweep over a 300× range in controlled synthesis, and audio-visual fusion under swapped audio and degraded vision on in-the-wild video. We evaluate three frontier model families against matched blind controls, a spectral oracle, linear probes on frozen audio encoders, and a human listener. One model is statistically indistinguishable from its blind control. The other two double-dissociate: each succeeds on the audio distribution where the other fails. No model succeeds at fine-grained material identification or size ordering, where a linear probe on a frozen 74M-parameter speech encoder outperforms every model. Under audio-visual conflict, models report the video’s material even when instructed to judge by sound. The relevant information is linearly accessible in generic audio embeddings; current models fail downstream of perception. We release item lists, the synthesis toolbox, and evaluation code.