ROAD: Rule-Grounded Context-Aware Open-World Driver Anomaly Detection
Abstract
Driver monitoring concerns determining not only what a driver is doing, but whether the behavior violates a driver-behavior safety rule under the current driving context. Existing distracted-driving benchmarks largely reduce this problem to closed-set action recognition, often missing cases where the same behavior changes meaning with road scene, vehicle state, duration, and rule-specific exceptions. We present ROAD, a rule-grounded context-aware open-world framework for driver anomaly detection. ROAD represents driver-monitoring constraints as structured rule cards and jointly reasons over in-cabin driver evidence, road-context video, optional vehicle-state metadata, and open-world semantic cues to produce structured anomaly judgments: driver action, violated rule, anomaly score, supporting evidence, exception status, and explanation. Unlike generic MLLM prompting, ROAD separates perception, rule grounding, and semantic generalization. Road context modulates driver-centric evidence, rule-card encodings select rule-relevant evidence, and a LoRA-adapted MLLM extracts open-vocabulary safety cues from compressed video evidence. This design targets two forms of open-world generalization: recognizing unseen driver behaviors beyond a fixed action taxonomy and adapting to new or revised rule cards. We further introduce ROAD-Bench, a rule-grounded benchmark with context, exception, evidence, anomaly-score, and explanation annotations. With a three-stage curriculum from action primitives to context-rule alignment and anomaly adaptation, ROAD outperforms evaluated baselines across driver-distraction recognition, open-world behavior recognition, and rule-grounded anomaly detection, using an 8B multimodal backbone. It achieves 92.01\% average accuracy over seven driver-recognition benchmarks, 77.9\% AUROC and 46.9\% unseen-class F1 in behavior recognition, and 73.6\% AUROC, 68.3\% rule accuracy, and 54.5\% Expl.-Rule F1 on ROAD-Bench.