Beyond Watching: Active Behavioral Monitoring of Black-Box AI Agents
Abstract
Monitoring AI agents is challenging when their internal reasoning is unavailable or cannot be relied upon. We propose active behavioural monitoring, a framework in which a monitor asks follow-up questions and observes whether an agent changes its answer to match what the questions suggest. We give agents tasks with impossible success conditions and examine cases where they reward hack or claim success despite failing the task. The monitor then asks questions that suggest a different answer or explanation, and tests whether the agent changes its claim to fit the monitor’s suggestion by editing existing evidence. The framework requires no access to model weights, activations, or private chain of thought. We propose evaluating it across multiple models in coding and data analysis tasks with independently verifiable outcomes. Our study compares passive observation with neutral and leading follow-up questions, including questions that suggest correct and incorrect explanations. We measure whether the resulting answer changes help identify failures that were difficult to detect from the original interaction alone. The central question is whether an agent’s willingness to reshape its claims around a monitor’s questions provides a useful signal for detecting unreliable behaviour.