Dose Blindness and Motivated Confabulation: A Multi-Layer Behavioral Audit of LLM Health-Impact Recommendations
Abstract
Large language models are increasingly deployed as decision-support agents in health-relevant domains, yet evaluation typically reduces to one question: did the model recommend the correct action? We introduce a three-layer audit that decomposes LLM reliability in pollution-aware urban routing into route selection, directional reasoning and numeric quantification, revealing a reliability gradient invisible to standard evaluation. Across four families (Mistral, Llama, Qwen, Gemma; n=498 origin-destination pairs each), route selection is correct in 49-96% of cases and directional reasoning in 83-88%, but numeric health-impact claims are accurate within 5 percentage points in only 34-55%. Quantification is the weakest layer in all four families; the full ordering from selection down to quantification holds in three of them. Three failure modes are universal: stereotype anchoring, with 30-48% of claims clustering in a fixed 10-15% band against an actual-value base rate of 15.1%, with a large point mass at exactly 15%; motivated confabulation, with signed overclaiming scaling with the concentration-dose gap (r=0.38-0.67, all p<0.001); and dose blindness, with three of four models recommending the lower-concentration route in 93-96% of cases without accounting for exposure dose. Injecting an explicit dose formula cuts hallucination by 13 percentage points and collapses the point mass at exactly 15% (19.1% to 1.6%), but the confabulation correlation survives at r≥0.45 under every scoring we tested. Whether the injection weakens it depends on the construct scored (p=0.69 to p=0.019 on the same responses), and neither reading is adequately powered. The failure is thus two-component: a prompt-addressable learned prior, and a residual scaling that instruction does not remove.