How Much Physics Can Vision Learn from Audio? Probing and Distilling Damping Structure
Zainab Afolabi ⋅ Vadim Selyutin ⋅ Zhao Gao ⋅ Babs Khalidson ⋅ Mohamed Ghoneim
Abstract
Physical understanding requires inferring properties like damping that govern material behavior, yet vision models struggle with such latent structure while audio encodes it directly through measurable acoustic decay. We probe frozen encoders for the quality factor $Q$, a dimensionless measure of resonance, and find that audio embeddings predict $Q$ at Spearman $\rho = 0.73$ while single-frame vision reaches only $0.28$. To transfer this physical structure, we distill a frozen audio teacher into a frozen CLIP encoder through a trained projection head. Under a leave-one-material-out protocol where category recognition earns nothing, distillation improves correlation from $0.234$ to $0.338$ (15/17 materials, $p=0.002$). Controls confirm the gain requires audio-visual correspondence rather than added capacity. Four frames spanning the impact provide half this improvement without any acoustic supervision, indicating that part of what single-frame vision misses is visible dynamics. The teacher reaches $0.600$ under the same protocol, so most of the acoustic signal does not cross.
Chat is not available.
Successful Page Load