Inter-Sensor Delay as a Learnable Physical Characteristic
Abstract
Self-supervised objectives learn to represent a physical system from its sensor measurements. Propagation and transduction displace the instant at which an event reaches each sensor and the representation space, hence co-indexed samples can describe distinct events. Multimodal objectives typically fix the inter-sensor delay at zero, replacing a physical characteristic of the sensing apparatus with an inductive bias. We propose LODES, a joint-embedding predictor that learns offset distributions on edges between streams without timing labels and predicts each stream from the other streams through the distributions, estimating a directed delay graph from their expectations. The recovered graph reproduces the order in which events reach the sensors. The representation matches or improves on synchrony-assuming baselines across seven benchmarks, and re-aligning a desynchronised stream at the recovered offset increases frozen-probe accuracy, which enables inter-sensor timing to be modelled and corrected without annotation.