RIFT: Towards a data-driven taxonomy of clinical chatbot safety risks
Abstract
Clinical conversational large language models are increasingly under regulatory and clinical scrutiny. While several benchmarks and expert-guided taxonomies have been produced to systematically map the risk space, there is no unifying taxonomy which is data-driven, actionable, and exhaustive. We present RIFT (Real-world/Research Integrated Failure Taxonomy), the first data-driven taxonomy characterising clinical chatbot risks across 22 risk dimensions. We built this by extracting 122 literature-derived failure modes from 111 sources, and factor analyse their co-occurrence across 14,232 health-related real world and benchmark conversations. Using this taxonomy, we descriptively analyse risk dimension coverage across published benchmarks, demonstrating that no benchmark covers the full risk space. We also show that the literature has redundancy in failure modes, where dimensions fit on 20 sources predict presence of unseen failure modes to saturation. More broadly, our work provides an iterative, data-grounded taxonomy of clinical risk domains enabling standardisation across future deployment-focused evaluations.