Position: Health Software Compliance Must Be Automated Before Agents Outpace Reviewers
Abstract
Health software is often subject to a dense network of legal obligations, and those obligations are increasingly discharged in code. Going forward, a growing share of that code will be written and updated by coding agents operating at machine speed. Because this will often remove the human review step at which code compliance was historically analyzed (and, if applicable, remedied), we take the position that, going forward, automated compliance will be a precondition for deploying agentic software development in health. We recognize, however, that automated compliance in this sector is especially challenging for two reasons. First, many legal obligations in health turn on criteria whose determinative evidence sits outside the codebase (e.g., in institutional records). Second, analyzing health software compliance often requires expert judgment, whose automation, while increasingly tractable with large language models, requires additional investment related to validation (e.g., expert-annotated benchmarks). In response to these hurdles, our central contribution is a decidability grid that classifies any health software legal obligation according to whether its determinative evidence lies in the codebase and whether expert judgment is required. Each of the four classes in our resulting grid corresponds to an investment (none, integration, validation, or both) standing in the way of automated compliance. For those building agentic health software, the grid is a way to triage efforts, tackling aspects of compliance that can responsibly be automated today (versus those that cannot). After introducing our decidability grid, we make a call to action: both to researchers, to build the validation infrastructure that makes expert judgment automatable, and to lawmakers, to migrate legal obligations toward codebase-decidability - with the goal, in each case, of helping health software improve at the rate other domains already do.