Comparing General and Failure-Mode Targeted Synthetic Datasets for Bounding-Box Object Detection
Abstract
Object detectors often fail on blurry, dark, and small objects. A common remedy is to augment scarce real training data with synthetic images, but most prior work does not distinguish between synthetic data that broadens coverage of the input distribution and synthetic data that targets the detector's known failure modes. We compare the two strategies under a fixed budget. Starting from a YOLOv12n detector fine-tuned on a merged plastic-bottle dataset, we train three models: one on real data only, one on real data plus 111 general synthetic images, and one on real data plus 111 synthetic images generated to depict three failure modes (blur, low-light, and small objects). Both synthetic sets are produced with SDXL and differ only in composition. We evaluate all three models on a shared clean test set and on corrupted versions of it, applying each failure mode at five severities. Neither synthetic strategy changes mean average precision under corruption by more than 0.03, a margin consistent with run-to-run variation, and the targeted set shows no reliable advantage over the general one. However, both augmented models raise recall under blur substantially (from 0.770 to 0.881 and 0.861), improving F1 while precision stays flat. The results suggest that at this budget, roughly 7\% of the real training set, the composition of synthetic data matters less than whether it is added at all, and that its clearest effect is on recall under degradation rather than on localization quality.