From Exploit Completion to Auditable Evidence: A Position on Closed-Loop Agentic Vulnerability Discovery
Abstract
Evaluations of AI security agents usually emphasize terminal outcomes such as captured flags, completed exploits, or solved tasks. These measures are useful, but professional vulnerability research also requires claims whose evidence, authorization, and provenance can be inspected. We describe Closed-Loop Agentic Discovery (CLAD), a reference protocol that separates hypothesis formation, risk authorization, controlled execution, evidence adjudication, and persistent memory. Each active claim must remain connected to protected observations or declared assumptions. Target-controlled output is treated as untrusted data, composed claims cannot acquire assurance through composition alone, and retrieval is restricted to evidence-qualified prototypes and counterexamples. Because operational findings may be confidential, the proposed evaluation relies on public vulnerable and secure testbeds, causal ablations, and claim-level scoring rather than undisclosed outcomes. An anonymized artifact schema, validator, and synthetic example accompany the paper. CLAD is presented as a falsifiable research agenda: exploit completion remains informative, but the object of evaluation is reproducible knowledge produced within an explicit authorization and risk boundary.