Evaluating evidence use in medical deep research for diagnostic decision support
Abstract
Medical deep research agents combine patient information with repeated search and synthesis, but access to relevant sources does not ensure a correct or well-supported diagnosis. We evaluate 200 multiple-choice and 69 open-ended multimodal diagnostic cases using four medical evidence sources. We compare no search, fixed retrieval, and agentic search, measuring diagnostic correctness, token use, and two automated report-error categories: feature-recognition errors and source-claim support errors. The highest observed accuracy is 49.5\% for multiple-choice cases and 24.6\% for open-ended cases. Search produces both corrections and regressions, and more elaborate workflows do not consistently improve accuracy. Recorded case studies illustrate failures to consider the correct diagnosis, select it despite relevant evidence, and preserve source attribution during final synthesis. These results support evaluating diagnosis correctness and report evidence separately, and examining how patient findings and retrieved sources are used throughout diagnostic research.