An Evaluation of Synchronization and Adaptive Routing Performance in LLM Multi-Agent Systems for Clinical Diagnosis
Abstract
The widespread adoption of Large Language Model Multi-Agent Systems (LLM-MAS) across high-stakes domains such as industrial management, drug discovery, web development, and scientific research has ushered in a wave of research examining LLM-MAS's viability in clinical diagnosis. We examined the routability and coordination potential for State-of-the-Art (SOTA) clinical LLM-MAS. The study analyzed the performance of three Large Language Model (LLM)-based routing approaches across three models and two medical datasets. Through the first systematic analysis of LLM-based routing among clinical MAS architectures, we find that the SOTA architectures have correlated correctness in all architecture pairs, that two architectures can solve almost every medical case, that less capable LLMs paradoxically complement each other, and that LLM-based routing usually fails to outperform a random router and fails to reach the accuracy ceiling of a perfect router. These results supplement the basis for further investigation into the quirks of both clinical and broad LLM-MAS as a whole, for the eventual purpose of identifying LLM properties that aid the medical field, while also discovering potential limitations in prior theories regarding clinical LLM-MAS.