Healthcare.pdf: Evaluating Deliverables, Not Answers, in Health Occupations
Abstract
Much professional work begins with a document: a guideline, a drug label, a clinical report. A professional’s deliverable is rarely judged on correct answers alone: every claim must be traceable to its source, and where the source runs out, the deliverable must say so. Benchmarks score the answers, but in healthcare an answer that cannot be attributed cannot be acted on. We introduce Healthcare.pdf, a healthcare benchmark of 179 tasks spanning three occupational personas, each a family of related healthcare occupations. Each task sets a piece of an occupation’s work on a document it uses and asks for one deliverable. That deliverable is scored against criteria at levels of generality, from those every document-grounded answer must satisfy to those particular to a role or a single task, so a failure can be attributed to a level rather than to the rubric as a whole. Across six frontier models, correct answers do not make acceptable deliverables: the strongest model reaches .799 task-strict correctness but only .184 full-rubric pass. Source traceability is highly instruction-sensitive: on a 55-task ablation, explicitly requesting citations raises GPT-5.6 Sol from .164 to 1.000 traceability pass with little change in correctness, yet full-rubric pass remains only .091. Thus, prompting can repair one professional requirement without closing the broader deliverable gap.