Learning the Temporal Offset Between Streams
Abstract
An interactive multimodal system perceives through multiple streams and responds before the streams are complete. Each stream reaches the model with its own latency, hence a shared event can be recorded at a different instant in each stream, and a latent state constructed from co-indexed samples pairs measurements of distinct events. Existing multimodal objectives fix every cross-stream offset at zero rather than exposing a coordinate for it. We propose OFFBEAT, for offset recovery by equivariant alignment of timelines, a joint-embedding predictor learning one distribution over candidate offsets and one directional influence for each ordered stream pair. Correlation over the candidates makes the reported offset follow a displacement applied at test time, with no timing label. Re-aligning a desynchronised stream at that offset restores frozen-probe accuracy, which enables an interactive system to correct the alignment of its own streams without annotation.