Verifying Laboratory-Agent Proposals with Policy-Conditioned Error Memory
Abstract
Laboratory-agent safety verifiers must distinguish hazardous experiments from legitimate research using similar reagents, equipment, and procedures. We study an offline, single-turn verifier that combines a 21-rule tiered policy with label supervised external error memory. Labelled calibration errors become natural language reflections retrieved from frozen memory at evaluation; model weights do not change. Tests cover a boundary-case split, policy and memory interventions, external benchmarks, and rule-labelled real compounds. Across three calibration orders on NearMiss500, memory raises F1 from 0.158 ± 0.049 to 0.757 ± 0.018 after five iterations. In a single-seed factorial evaluation, restoring explicit policy text alongside policy-conditioned memory reduces FPR from 0.246 to 0.108, while TPR falls from 0.743 to 0.600 and F1 from 0.675 to 0.667.A supervised classifier remains stronger (F1=0.941), and external tests expose sensitivity to task representation and provider filtering. The evidence supports in spectable, weight-free adaptation of offline proposal decisions; it does not isolate policy-free memory or establish safe autonomous operation.