A Comprehensive Framework to Understanding Generative AI Performance in Healthcare
Abstract
Existing approaches to evaluating clinical AI largely focus on response accuracy on exam-style benchmarks, which often fail to capture the uncertainty and complexity inherent in real-world clinical decision-making. As a result, strong benchmark performance may not reliably translate into effective performance in realistic clinical scenarios. To address this limitation, we present a modular, multidimensional framework for evaluating generative AI along five complementary dimensions: correctness, safety, reasoning quality, uncertainty and calibration, and clinical anthropomorphism. Our framework decouples data ingestion, output formatting, and evaluation, enabling straightforward integration of new models, datasets, and evaluation metrics. We use this framework to assess four open-weight models: Gemma-4 26B, Gemma-4 2B, Llama3 8B, and the domain-specific MedGemma 4B, across three datasets spanning increasing levels of clinical complexity: MetaMedQA, DiReCT, and ER-REASON. This multidimensional evaluation reveals important failure modes that conventional accuracy measures can overlook. In particular, fluent, convincing responses can conceal deficiencies in safety and reasoning, while models show substantial miscalibration and overconfidence when confronted with open-ended diagnostic problems. These findings underscore the limitations of accuracy-centric evaluation for clinical AI and position the framework as an extensible tool for rigorous, safety-oriented assessment of generative models in clinical settings.