Cross-View Representation Alignment for Robust Multi-View Driver Action Recognition
Abstract
Driver distraction is a leading cause of traffic accidents, making accurate driver action recognition important for road safety. A single camera view often gives an incomplete picture, missing actions due to occlusion or unfavorable viewing angles. Using multiple camera views can help, but introduces a less obvious problem: because each camera observes the same action from a different geometric viewpoint, the resulting representations do not naturally align even when they describe the same physical action. Most fusion methods overlook this geometric discrepancy and combine views directly, forcing the fusion step to both correct for mismatched representations and exchange useful information at once. We propose Janus, a dual-view framework that explicitly separates these two steps. Representations from the two views are first aligned using supervised contrastive learning, then combined using bidirectional cross-attention. On the DMD benchmark, Janus achieves 93.58\% macro F1 and 98.53\% top-1 accuracy, outperforming existing multi-view methods. The learned alignment substantially improves cross-view retrieval accuracy R@1, from 0.05–0.12 to 0.98, indicating that representations from different views become strongly aligned before fusion despite their differing physical viewpoints. Janus also demonstrates improved robustness under zero-shot occlusion, suggesting that explicitly aligning complementary views before information exchange may contribute to improved robustness.