How Much Physics Can Vision Learn from Audio? Probing and Distilling Damping Structure
Zainab Afolabi ⋅ Vadim Selyutin ⋅ Zhao Gao ⋅ Babs Khalidson ⋅ Mohamed Ghoneim
Abstract
Vision models struggle to infer properties like damping that govern how materials behave. Audio encodes such properties directly, because an object's acoustic response to impact reveals its damping through measurable decay. We probe frozen encoders for the quality factor $Q$, a dimensionless measure of how many oscillations an object completes before its ring dies away, and report how well each embedding orders impacts by $Q$ using the Spearman rank correlation $\rho$, for which $\rho = 1$ is a perfect ordering and $\rho = 0$ an unrelated one. Audio embeddings reach $\rho = 0.73$ and retain most of that advantage on materials withheld from training, while a single video frame taken at the moment of impact supports only $\rho = 0.28$. We transfer this representation by distilling a frozen audio teacher into a frozen CLIP backbone through a trained projection head, and find that this recovers a useful fraction of the audio advantage. Because pooled correlations on this task are bounded in practice by category identity, we also evaluate under a leave-one-material-out protocol in which the held-out material is withheld from the projection head as well as the probe. Distillation improves the correlation there from 0.234 to 0.338 on 15 of 17 materials, a margin no category-level strategy can produce. Widening the visual input to four frames spanning the impact supplies half of that improvement at no acoustic cost, which indicates that part of what a single frame misses is visible dynamics.
Chat is not available.
Successful Page Load