StomataBench: Measuring Taxonomic Generalization in Stomatal Detection
Abstract
Accelerating stomata measurement is essential for developing drought-resilient plants. However, most automated stomata detectors are tested under narrow imaging and species conditions. Their behavior under broader ecological distribution shift remains unclear. We introduce StomataBench, a benchmark for stomatal object detection across species, acquisition protocols, and biological density regimes. The benchmark uses a curated multispecies training set of 3,154 images with 134,802 annotated stomata from 21 woody species. It is evaluated on four test datasets totaling 19,657 images, covering diverse woody taxa and heterogeneous public crop sources. Unlike standard detection benchmarks that report only AP, StomataBench evaluates measurement reliability. We measure COCO AP, F1, score-threshold sensitivity, image-level count error, and stomatal density agreement. We benchmark sixteen detectors spanning two-stage, one-stage, transformer, open-vocabulary, and promptable foundation models. The results show that model rankings change with the biological objective. GDINO-SB provides the strongest average AP generalization, while SAM3-OD gives reliable operating point for count and density estimation. Conventional detectors remain competitive on near-domain woody plants, but foundation models are robust under harder ecological shifts. Failure analysis shows that far-OOD detection is often limited by missing stomata. These results identify cross-lineage proposal generation as a key unsolved problem for reliable stomatal phenotyping under ecological distribution shift.