Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants
Abstract
Refusal rates conflate upstream provider behavior with downstream safeguards and do not reveal what information remains in a delivered answer. Safeguard-conditioned uplift is an evidence-gated framework for evaluating deployed access paths for dual-use biology assistants. It attributes interventions to their source, selects thresholds under a legitimate-access budget, and supports content claims only when answer verifiers pass frozen qualification tests. In a prospectively frozen route and generation study, two of six configurations pass the action-level rule, but both inherit the same Claude Opus provider effect rather than independent downstream gains. On 104 unused matched pairs, selectivity transfers, but benign intervention exceeds the 20\% limit. No joint-scoring verifier meets the frozen criteria; a factorial follow-up shows that criterion isolation and ordinal outputs improve edit-direction accuracy. The evidence supports action-level selectivity and verifier interface effects, but not calibrated access, verified content removal, or biological-risk reduction.