Silent Failures in Agent Oversight: A Null-Supervision Control for Cross-Model Code Critique
Abstract
Agent oversight increasingly relies on one agent - a critic, a safety monitor, a reward model - to supervise another and intervene when something looks wrong. This is subtler to evaluate than it looks: comparing "with oversight" to "no oversight at all" confounds the informational content of the supervisory signal with the mere act of intervening, since a revision step can introduce or remove errors regardless of what the overseer actually said. We introduce a null-supervision control - an evaluation arm that performs the intervention exactly as in the supervised condition but empties it of content - and use it to evaluate a minimal oversight system: a critic agent that reviews and a reviser agent that revises a subject agent's code across multiple rounds. Measured against this control, oversight is not merely unhelpful, it is actively harmful: a cross-model critic reports a bug in already-correct code on 0.945 of opportunities (1106/1170), driving a significant drop in task accuracy relative to the null-supervision floor. The failure is largely invisible to activity-based monitoring - the kind of monitoring a deployer without a ground-truth oracle would actually have: the loop rewrites up to 0.317 of outputs in a single round while genuine repair never exceeds 0.028, so a monitor watching only whether the overseer is "doing something" would observe engagement, not damage. Self-critique, in contrast, collapses into the null control itself: the critic emits no usable signal, so it neither helps nor hurts. Adding an "expert" persona to the cross-model critic raises its false-alarm rate further (to 0.989, consistently across all three seeds) and suppresses genuine repair, though the net effect on task accuracy is not statistically separable at this sample size. We argue the null-supervision control is a general primitive for evaluating agent oversight: without it, a deployer cannot attribute a change in a supervised agent's behavior to the supervisory signal's content rather than to the cost of oversight's mere existence - exactly the blind spot that lets an unsafe overseer look like a working one.