Are Agent-Safety Benchmarks Ready for Monitor-Portfolio Procurement? Integrity Protocols and Data Feasibility Limits
Abstract
Selecting a fixed subset of heterogeneous safety monitors under monetary, token, or latency constraints is a concrete deployment problem. Before comparing portfolios, however, a measurement contract must bind task-level units, benign controls, missing-output semantics, and a selection lock that predates held-out evaluation. We design a template-aware, fail-open procurement protocol and apply it as a readiness audit to two public agent-safety resources. The audit yields three findings. First, the official InjecAgent archive has 63,240 execution rows across 60 model-prompt matrices, but these repeat 1,054 crossed cases formed from only 17 distinct source user templates and provide no matched clean executions. Second, same-code parser replay agrees with released binary states on 63,225 rows, while 18,722 are marked invalid and 663 unparsed; this establishes parser reproducibility, not independent label validity. Third, current releases fail the measurement contract for confirmatory procurement: the intact template graph cannot yield a three-way split, natural benign controls are absent, and a template-disjoint 5/7/5 matched pilot leaves only five procedural calibration controls versus 404 independent benign templates required by the frozen precision target. A bounded, post-hoc paired intervention over 17 source clusters observes 22.1 percentage points fewer explicit target-tool proposals in the empty-placeholder arm (source-cluster bootstrap interval 11.8–33.8); a remotely witnessed follow-up has the same direction but reuses those source clusters. No benchmark tool is executed. Exact-oracle simulation is not evidence that a portfolio is better on real agent data: under a stylized generator and fixed benign calibration, it shows that adding harmful templates reduces ranking error while worst-cell oracle screening remains above 90%. We therefore report a measurement-readiness determination, not a confirmatory portfolio comparison. Current releases are row-rich but template-poor; the intended hypotheses remain deferred pending new source tasks and natural benign controls.