Reference-Based Directed Granger Gain for Evaluating Speech-Conditioned Listener Motion
Abstract
The bidirectional nature of conversation requires interactive avatars to take the paired speaker's audio into account when generating listener motion, but conventional motion-based metrics cannot establish whether that audio contributed to the result. We introduce the Reference-Based Directed Granger Gain (R-DGG), which measures the additional predictive information carried by speaker speech after accounting for listener history and speaker motion. We evaluate R-DGG on ground-truth listener motion, three dyadic systems, four talking-head generators, and mismatched speaker-listener pairs. Conventional metrics do not consistently distinguish methods with access to paired audio from those without it, whereas 95% R-DGG intervals lie above zero for ground truth and all dyadic systems and include zero for every negative control. R-DGG thus provides a measure of speech-conditioned listener motion that complements existing motion-based evaluations. We release the R-DGG implementation and precomputed EMOCA and LivePortrait features to support reproducibility.