Object Detection Benchmarks are Incomplete: The Role of Label Errors and Annotation Uncertainty
Abstract
While object detection has advanced through improved architectures and open-vocabulary models, we show that benchmark quality is fundamentally limited by annotation incompleteness. Across four widely used datasets (COCO, Pascal VOC, Cityscapes, KITTI), re-annotation reveals substantial increases in annotated objects (e.g., up to +60% on KITTI and +40% on COCO), driven primarily by previously unlabeled small, occluded, or densely packed instances. While some differences arise from dataset-specific annotation conventions, we consistently find that missing annotations are the main source of label errors across all datasets. To achieve high data quality, we introduce a scalable annotation pipeline that emphasizes high recall and captures ambiguity through soft labels aggregated from at least 11 annotators per object. The resulting annotations improve coverage and align well with human calibration. We show that benchmark performance is highly sensitive to annotation quality, although model rankings remain largely stable. We introduce two large-scale benchmarks: (i) an uncertainty-aware object detection benchmark, and (ii) a label error detection benchmark grounded in real label errors. We demonstrate that current detectors are misaligned with human uncertainty, and that training with soft labels improves calibration and better reflects human perception. Existing label error detection methods, which have been shown to perform well on synthetic noise, struggle to achieve high recall and precision on real-world label errors. Our results highlight the need for future object detection benchmarks to move beyond deterministic annotations toward high-recall, uncertainty-aware evaluation.