The Safety Tax: Quantifying the Over-Refusal Cost of Explicit Safety Instructions in Clinical FHIR-Write Agents
Abstract
A clinical LLM agent that writes to a patient record faces a trade-off between unsafe execution and wrongly refusing a safe order, yet no benchmark measures it with matched pairs and judge-free ground truth. We introduce MedAgentRefuse, a 100-pair matched safe/unsafe FHIR-write benchmark with deterministic, FHIR-state-based ground truth, and use it to evaluate four LLM agents under three prompt conditions (default, hazard-specific, and taxonomy-neutral generic safety instructions) across three seeds, for 7,200 primary task-instances. Under the default condition, harmful execution stays near floor for every model (rate at most 1%, 95% Wilson upper bound at most 2.4%). Under a paired task-cluster bootstrap, the hazard-specific instruction raises over-refusal on matched safe controls by 5.0-41.7 percentage points, and the generic instruction by 37.3-74.0 points; both are distinguishable from zero for all four models, and significantly larger under the generic instruction in every case. Vocabulary overlap between the hazard-specific clause and the benchmark's hazard taxonomy is not necessary to produce the effect, though the two instructions also differ in evidentiary threshold, so vocabulary is not isolated as the cause. Neither instruction produces a statistically resolvable reduction in harmful execution below the near-floor default rate, so the safety tax is both model- and wording-dependent. All results come from a synthetic, single-institution FHIR sandbox and are not a clinical-safety or deployment-readiness claim.