Removing Model-Level Risk Does Not Remove User-Level Risk in a Deployed Health Application
Kritika Chugh
Abstract
Work on trustworthy generation assumes that making a model safer makes the people who rely on it safer. We report a deployed case where that translation fails, and we measure where it fails. Our label-reading application, which answers questions about packaged food for people managing food allergy, enforces a guarantee of a kind that instruction-level alignment does not provide: its generative models may transcribe, select and narrate, but no model may compute a quantity, so fabricated derived values are excluded by construction rather than discouraged by training. Yet on a released 200-product corpus, users of the deployed system still face an 8.7% risk per scan (95% CrI [4.9, 14.1]) that a true allergen fails to surface, and that risk is structural: it depends on which of two paths a scan takes, not on anything a model invents. Three findings generalize. First, the metrics that certify models mismeasure user harm: faithfulness to a retrieved record scores an answer from an incomplete source as correct, and the gap between token-level accuracy and verdict-level harm bends in opposite directions on the two paths, overstating harm where records disagree and understating it where extraction fails. Second, failures compose: extraction fails at whole-panel granularity, so redundant declaration offers no protection, and observed harm exceeds the prediction of an independence model ($p = 3.7 \times 10^{-4}$). Third, the failures that matter are silent: no inexpensive signal separates a failed read from a clean one, and 48 blank results found in telemetry had produced zero user reports, so post-deployment monitoring must be instrumented rather than complaint-driven. We close with three practices: safety claims stated as auditable invariants whose violations are disclosed, evidence standards that report intervals and decline what the data cannot support, and monitoring designed for silence.
Chat is not available.
Successful Page Load