Oversight Under Flooding: An Unidentifiable Attack and a Robust Defense
Debarchan Basu
Abstract
An adversary need not fool a verifier. Instead, it can change the conditions under which the verifier operates. Deployed agent architectures terminate in a human approving consequential actions, and that human is modelled throughout the literature as an oracle: fixed accuracy, unbounded throughput. We treat the approver as a capacity-constrained component and ask what an adversary achieves by injecting benign filler–perturbing no individual decision, changing only volume. Our central result is a comparative static: the sign of an optimising defender’s threshold adjustment under flooding is governed by $\eta'$, the derivative of the elasticity of reviewer reliability with respect to load, and the slope of the reliability curve does not appear, though it is the quantity practitioners reason about. Past a critical volume $a_{crit}$ the optimising defender’s threshold jumps discontinuously off the selective branch to a saturated one–and we prove this happens whenever a saturated reviewer retains nonzero reliability, which makes the queue’s overflow discipline the parameter that selects the failure mode. Neither branch preserves oversight: a queue that sheds load never bifurcates but has no reliability floor either, and hides its own collapse behind a reviewer who looks fully utilised. The practical consequence is negative: on a guard whose ROC we measure from 369 real agent-injection pairs (AUROC 0.864, 95% CI [0.800,0.922]), $a_{crit}$ ranges from zero to twenty-eight times baseline traffic across plausible reviewer fatigue curves–zero meaning the defender is already past the jump with no adversary present–and by a factor of 98 across the guard’s own confidence interval alone–so we cannot distinguish a world in which one principal within its normal budget suffices from one requiring fifty coordinated identities. Against that, we prove that a defender conditioning its harm prior on observed load weakly dominates a load-naive one on expected harm, under a single explicit assumption and with no regularity conditions; a 46,875-configuration sweep finds the median gain is 10.4% and that it is negative nowhere. The asymmetry is the contribution: the attack’s cost cannot presently be stated, but a defense against it can be recommended without waiting for the missing measurement.
Chat is not available.
Successful Page Load