Static Benchmarks Are Not Enough: The Need for Dynamic, Adversarial Evaluation of PHI-Handling Medical Agents
Abstract
Security evaluation for medical generative AI still relies on static benchmarks: fixed inputs, fixed environments, non-adaptive attacks, and aggregate scores. These tests support regression and comparison, but they provide weak evidence about deployed, tool-using medical LLM agents that process protected health in5 formation (PHI), ingest attacker-influenced content, call tools, and apply role- and context-dependent authorization rules. Failures in such systems arise through retrieval, identity ambiguity, workflow misuse, authorization bypass, indirect prompt injection, and multi-turn elicitation. We argue that PHI-handling medical agents require dynamic, adversarial, deployment-specific benchmarking. Static benchmarks under-cover adaptive attacks, deployment-specific tool surfaces, and benchmark contamination. HIPAA, GDPR, the EU AI Act, and FDA lifecycle oversight do not prescribe adaptive benchmarking, but they do require continuing evidence that safeguards remain effective. We propose a tiered security-evaluation framework that retains static benchmarks as a baseline and applies adaptive evaluation to higher-risk PHI deployments. Our position is that, for PHI-handling medical agents, static benchmarks are necessary but structurally insufficient evidence of real-world security, and that dynamic, adversarial, deployment-specific evaluation should become part of the standard assurance process.