Critical-Timestep Observation Attacks in Multi-Agent Reinforcement Learning for Autonomous Driving
Abstract
Random perturbation tests may underestimate the vulnerability of autonomous-driving policies because identical perception errors can have very different effects depending on when and where they occur. We study a constrained observation-only attack against multi-agent deep Q-networks: one nearby-vehicle row is replaced by its value from the preceding timestep, while simulator state, rewards, actions, and policy parameters remain unchanged. A driving-risk score adaptively selects the attack timestep, victim agent, and observation row. Across ten independently trained policies in both Highway and Merge, critical targeting causes substantially greater return and crash damage than fixed-target random stale-row corruption. Matched replay on 800 Highway states further shows that the highest-scored stale-row intervention causes greater downstream harm than a random alternative from the same state. Robustness-oriented training improves random-corruption robustness but does not eliminate adaptive vulnerability.