Not All Instability Is Dangerous: A Type-Axis Decomposition of Clinical-Agent Divergence on Real Patient Records
Abstract
As large language models (LLMs) are increasingly deployed as clinical decision-making agents, evaluating their reliability has become increasingly important. However, existing studies quantify the amount of run-to-run variability without distinguishing whether that variability is clinically meaningful. We propose a type-axis decomposition framework that separates recommendations into six mechanistically distinct types by considering data-extraction errors, reasoning errors, and benign variation. The framework combines deterministic rule–based classification with targeted LLM adjudication in order to evaluate cases that require semantic interpretation. We evaluated the framework on 180 real NHANES patients across three frontier LLMs (Claude Sonnet 4.6, Gemini 2.5 Pro, and DeepSeek-V4-Pro) and five USPSTF preventive care recommendations, producing a total of 2,700 patient-recommendation evaluations. PATH divergence was the most common classification (54.6%). Of these cases, an LLM judge (GPT-5.4) found that 98.9% were paraphrases of reasoning and only 1.1% (16/1,474; 95% CI 0.6-1.8%) were clinically meaningful reasoning differences. Additionally, agreement across vendors was low (Cohen's κ = 0.13-0.42), indicating that models differ not only in overall accuracy but also in their failure modes. Meaningful divergence was especially shown in recommendations that showed heavy ambiguity within guidelines, such as race-specific thresholds, while simple criteria, such as age-based thresholds, produced no meaningful divergence. These findings suggest that treating all output variability as equivalent obscures clinically important differences. The proposed taxonomy provides a practical framework for evaluating clinical AI systems by distinguishing benign variation from instability that may require mitigation or human oversight.