Who Called? V33DA: A Physically Verified Multimodal Benchmark for Vocal Attribution in Zebra Finch Groups
Maris Basha ⋅ Yuhang Wang ⋅ Xiaoran Chen ⋅ Longbiao Cheng ⋅ Luca N Yapura ⋅ Anja T Zai ⋅ Mathieu Salzmann ⋅ Richard Hahnloser
Abstract
Deciding which member of a group produced a vocalization and where that vocalization was produced are central to studying social communication but are rarely evaluated with physically verified ground truth. Existing resources either localize sounds without identifying the caller, or recognize individuals from isolated vocalizations without reliance on the 3d candidate geometry. We close this gap with V33DA, a multimodal benchmark for spatial vocal attribution in social zebra finches: it comprises 33,625 verified events coupling 5-channel audio with three camera views, calibrated 3D pose for every visible candidate, and per-bird radio telemetry across 10 individuals and three experiments. Caller identity is verified from an on-body accelerometer-derived vibration signal; this channel is withheld from benchmark models and used only by an oracle ceiling. We also provide V33DA++ as an auxiliary extension with verified overlapping-call events and $\pm 2$s context windows, intended to support future source-separation and vocal-activity-detection studies. V33DA separates familiar-individual recognition from transferable spatial attribution. We evaluate candidate-conditioned caller attribution, with 3D source localization as a secondary diagnostic, under three regimes: session-disjoint testing, held-out-experiment transfer, and leave-one-individual-out evaluation. A broad set of reference methods exposes a consistent gap between in-domain accuracy and transfer under identity shift: methods that can exploit familiar vocal identity perform well on known callers but degrade sharply on unseen ones, whereas candidate-aware methods preserving explicit spatial reasoning transfer substantially better. The accelerometer oracle stays near-perfect across regimes, confirming the internal consistency of the accelerometer-verified labels and quantifying the ceiling available when the withheld verification channel is observed. We release V33DA with fixed attribution/localization protocols and V33DA++ with overlap and long-context tasks, together with code for the evaluated methods.
Chat is not available.
Successful Page Load