Reward Is Not Enough: Argmax-Induced Policy Collapse and Idle-Time Hazards in Autonomous Cyber Defense
Vasanth Iyer ⋅ Atharva Umbre ⋅ Vir Phoha
Abstract
Verifying an agent means answering a narrow question: when we change its policy, tools, or training procedure, did it actually get better? For reinforcement-learning (RL) agents that take real actions in an environment, the default verification signal is scalar episode reward, and we show on a network-defense RL benchmark that this default can be actively misleading. We compare PPO, LSTM-PPO, and DQN Blue-team (defender) agents in the CybORG Autonomous Cyber Operations gym and find that an argmax-induced, deterministically collapsed PPO policy -- one that repeats a single fixed action regardless of observed network state -- attains a substantially better raw reward ($-20.0$) than a DQN policy that visibly differentiates its behavior across states (six distinct action--host combinations; 39 action changes per 50-step episode; reward $-228.1$). We argue this is not classical train/test overfitting, for which no held-out adversary split exists here, but a \emph{reward--behavior divergence}: reward computed against one fixed, deterministic adversary trajectory certifies policies that happen to intercept that one path, not policies that condition on state, which is the property autonomous defense actually requires. We report a second, independent verification failure after retraining the DQN approach for a 20-host, multi-adversary environment: the resulting policy reduces mean penalty by approximately 59\% relative to random across the two active-adversary evaluation profiles, yet incurs roughly an order of magnitude \emph{more} penalty than a random policy when no adversary is present ($-87.5$ vs.\ $-8.7$) -- an idle-time availability hazard invisible to any evaluation that only measures performance under attack, and one that persists despite adversary-free episodes having been part of training. We connect both findings to the reward-hacking and specification-gaming literature, and present a STIX 2.1/MITRE ATT\&CK export pipeline with embedded per-technique Q-value audit metadata as one concrete instantiation of the heterogeneous, environment-grounded verification signals that we argue must accompany scalar reward for RL-based autonomous agents. Because each bundle is machine-readable structured data rather than free-text narrative, it also serves as a grounding substrate for LLM-based reasoning over agent behavior: technique sightings, Q-value preferences, and executed actions can be retrieved and cited directly, rather than inferred or hallucinated by a language model reasoning about the policy after the fact.
Chat is not available.
Successful Page Load