Retrieval-Augmented Verification of Chest X-ray Radiology Reports using COT and Consensus Reasoning in an LLM Judge
Abstract
With the emergence of large-scale vision language foundation models (VLM), radiology report generation has become a popular application area with realistic-looking reports being now generated. However, hallucinations are still prevalent in these reports making a verification step necessary before such reports could be viewed by clinicians. Recently, chain-of-thought (COT) reasoning vision-language judge models are emerging that can do reference-free evaluation but they are prone to self-preference bias and prefer hallucinated plausibility to clinical factuality when no objective ground truth is present. In this paper, we take a new approach to fact-checking by training such judge models to exploit the consensus present in prior radiology reports of similar patient images. Specifically, a new loss function is designed to capture the consensus of findings present in similar patient reports while allowing for patient-specific individuality and sparsity in the reporting of findings. We show through extensive studies across datasets, judge models and report generation models that such retrieval-based consensus-trained judge models improve the overall quality of automated reports by 8-17\% as measured through comparison with ground truth reports on the reported findings in evaluation settings.