Reinforcement Learning with Stochastic Substitute Control Under Infrastructure Failures
Salim Oyinlola ⋅ Barika-Bodunrin Y Olamide
Abstract
Reinforcement learning assumes that once an agent selects an action, the environment executes it. That assumption is silently underwritten by reliable infrastructure and fails wherever the channel carrying a control decision to the world is itself unreliable. We study this concretely via reinforcement learning for traffic-signal control under stochastic power outages, on a five-junction corridor reconstructed from OpenStreetMap geometry in Lagos, Nigeria, with demand calibrated to published measurements. We train a deep Q-network that outperforms fixed-time, actuated and max-pressure control, then evaluate it under an outage process governed by three stylised models of the substitution mechanism that takes over while the signal is dark. In this simulation, the sign and magnitude of outage cost depend strongly on the substitution regime: self-organised priority improves episode return by up to 57.9\%, while the modelled human-direction regime degrades it by up to 648\%. The sign persists across a nine-point sensitivity sweep ($+157\%$ to $+1167\%$) and across non-learned controllers, although these point estimates use five episodes per cell and should be read as descriptive. Every OpenStreetMap-signalised node in the study region also controlled a single approach, leaving adaptive control undefined without reconstruction. Traffic signals are the case study. Our contribution is to make an explicit, stateful substitute policy a first-class object alongside the existing literature on degraded actuation and communication dropouts.
Chat is not available.
Successful Page Load