Can World Models Play the Piano? Benchmarking Physical Consistency in Audio-Visual Generation
Abstract
Visual plausibility alone is insufficient for evaluating world models: generated interactions must also preserve the physical relationship between actions and their sensory consequences. A generated piano performance, for example, can look and sound convincing while the visible key presses and audible notes disagree in identity or timing. We introduce PianoBench, a prompt-defined benchmark that probes physical consistency in audio-visual generation through controlled piano performances. The piano's explicit key--pitch mapping links visible actions to acoustic outcomes, allowing both to be evaluated against symbolic event targets without reference performance videos. PianoBench spans five task levels: single notes, ascending sequences, descending sequences, repeated notes, and chords, testing event identity, count, order, repetition, and simultaneity under shared scene constraints. Our evaluation separately assesses visual action accuracy, audio accuracy, and audio-visual temporal alignment to distinguish errors in execution, acoustic content, and synchronization. An initial study of seven generation models reveals task-dependent weaknesses, with high automated visual scores sometimes accompanied by inaccurate note events and weak detected synchronization. PianoBench thus complements perceptual evaluation with an event-level diagnostic of whether generated actions and sounds remain consistent with a single physical interaction.