A Comparative Study in Surgical AI: Potential and Limitations of Data, Compute, and Scaling
Abstract
Background. Recent Artificial Intelligence (AI) models match or exceed human experts on many biomedical benchmarks, yet surgery—which requires visual recognition beyond text question-answering—is largely absent from prominent medical benchmark suites. Whether scaling general-purpose models suffices to deliver capable surgical AI, or whether specialized data is the binding constraint, remains unknown. Methods. We studied surgical tool detection on SDSC-EEA, a dataset of 67,634 annotated frames from 66 endoscopic endonasal neurosurgical procedures (20,016 held-out validation frames, split by procedure). We evaluated 26 models—20 open-weight Vision Language Models (VLMs; 2B–235B parameters) and 6 frontier proprietary models—under zero-shot prompting and LoRA fine-tuning, swept LoRA rank from 2 to 1024, and compared a 26-million-parameter specialized object detector (YOLOv12-m). Every protocol was replicated on three public datasets (CholecT50, PitVis-2023, SurgVU). Performance is reported as exact-match accuracy and micro-averaged F1. Results. Among 20 open-weight VLMs evaluated zero-shot on 20,016 held-out frames, none surpassed a majority-class baseline of 47.3% micro-averaged F1 (range, 1.0% to 44.9%); the highest general-benchmark scorer reached only 39.2%. The 26-million-parameter specialized detector (YOLOv12-m) achieved 54.7% exact-match accuracy (95% CI, 54.0 to 55.4) and 80.5% micro-F1—exceeding every zero-shot VLM by more than 40 percentage points—and matched a fine-tuned 27B-parameter VLM (51.1% exact match [95% CI, 50.4 to 51.8]; 81.4% micro-F1) using approximately 1,000-fold fewer parameters. Increasing LoRA adapter capacity by three orders of magnitude raised training micro-F1 from 66.3% to 99.5% but validation micro-F1 only to 69.9%, a persistent train–validation gap of roughly 30 percentage points. Across the three public datasets, the specialist and the fine-tuned VLM led every dataset (75.3% to 92.8% micro-F1), whereas all six frontier proprietary VLMs trailed both. Conclusions. Current general-purpose models face substantial obstacles in surgical perception that scaling alone does not overcome; the dominant bottleneck appears to be specialized data rather than model size or compute.