Fresh Flags Are Not Enough: A Retrospective Audit of an Agent Exploitation Task
Abstract
Fresh random flags make cybersecurity task outputs easy to grade, but do not establish how an output was obtained or why an agent stopped. We retrospectively audit existing artifacts for one hardened remote-exploitation task prepared for the Agents' Last Exam framework. A calibration summary counts 152 tool invocations; recovering its child trace increases the total to 264. Both traces end in context-limit API errors; the coordinator subsequently records zero successes in three grading rounds for the unfinished artifact. Later validation verdicts record reference success in three of three rounds and filesystem-probe nonrecovery in two of two. Source inspection finds that the grader's identity separation depends on launcher privilege, a condition absent from its verdicts. These observations support a bounded evaluation-validity result, not an estimate of model capability or a demonstrated sandbox escape. We propose reporting task outcomes, execution conditions, delegated work, and termination causes separately, with a minimal execution contract and controls that test their own ability to detect violations.