Healthcare LLM Benchmarks Are Only as Good as Their Explicit Assumptions
Abstract
Benchmarks are necessary for healthcare evaluation, but are not sufficient for predicting deployment performance. Our position is that the evaluation--deployment gap arises from implicit assumptions connecting benchmark conditions to how users interact with models in practice. We distinguish between task assumptions, which can be tested using interaction data, and outcome assumptions, which depend on human behavior and require behavioral or outcome data to test. By retrospectively analyzing a healthcare randomized controlled trial, we illustrate how the observed evaluation--deployment gap separates into task and outcome gaps of roughly equal size. To make these assumptions explicit and testable, we propose BenchmarkCards, an artifact for documenting assumptions connecting evaluation to deployment, and staged evaluation, a procedure for progressively testing these assumptions before deployment.