RETR: A Structure-Preserving RGB-Event Transformer for Robust 3D Lane Detection
Abstract
Robust 3D lane detection requires accurate metric road geometry recovery from visual inputs, yet conventional RGB-based approaches suffer from fragile lane evidence under adverse illumination, motion blur, and long-range perspective compression. Event cameras offer complementary high-temporal-resolution, high-dynamic-range structural cues that are robust to these challenges, but their sparse, motion-dependent responses cannot be directly fused into consistent 3D geometry. In this paper, we present the first attempt to introduce event cameras into 3D lane detection. To enable systematic research, we first build two multimodal benchmarks with metric 3D lane annotations: DSEC-3DLD (real-world sequences) and Ev-OpenLane (large-scale simulated sequences). We further propose RETR, a structure preserving transformer that converts complementary RGB and event observations into coherent 3D lane geometry. RETR first aligns modality consistent evidence through Reciprocal Context Flow Fusion, then preserves thin and uncertain lane structures with Uncertainty Aware Structural Consolidation, and finally decodes ordered lane hypotheses using a Geometry State Decoder with proposal conditioned initialization, reference conditioned query evolution, and diversified geometric inquiry. Extensive experiments show that RETR achieves state-of-the-art performance on both benchmarks, especially showing strong improvements in challenging lighting and complex road geometry scenarios.