Risk Without Boxes: Task Grounding Elicits Hazard Localization in Driving-Video World Representations
Abstract
End-to-end collision anticipation reaches state-of-the-art performance trained on nothing but a clip-level label -- whether the ego vehicle collides -- with no box, mask, or spatial target at any stage, yet read through its attention, the deployed model concentrates on the specific vehicle about to be involved. We dissect where this comes from across five variants of one lineage, evaluated out-of-distribution under the system's pre-specified late-block attention readout. The pretraining stage is a predictive world model in the joint-embedding sense, and on its own it does not ground the representation in what matters for safety: under all seven readouts evaluated it produces no consistent hazard localization. Collision supervision alone localizes strongly but shallowly and unstably, in early layers the deployed readout never sees; pretraining consolidates this into stable, reproducible binding in the late layers the system reads. Task accuracy does not reveal this: a variant that classifies well above chance points at chance under the production readout, and a matched multi-seed comparison shows distillation refining localization while leaving accuracy essentially unchanged. The attention-selected regions carry prediction-relevant evidence: deleting them collapses the prediction far faster than deleting the same patches in random order, on nearly every clip.