Taming the Long Tail in Federated Crop Disease Diagnosis: Adaptive-Threshold Semi-Supervised Learning Under Extreme Non-IID Heterogeneity
Abstract
Diagnostic AI for agriculture is often built and validated on conditions that do not hold in the regions that need it most: centralized data, balanced classes, abundant labels. Smallholder and regionally specialized farms cannot pool diagnostic images onto one server, whether for connectivity, cost, or data-sovereignty reasons, and each typically grows only a handful of crops, so any farm-level model has to work with scarce labels and a narrow, skewed view of the world. Federated learning (FL) avoids pooling raw data, but federated semi-supervised learning (FSSL) methods inherit a fixed pseudo-labeling threshold from centralized, class-balanced benchmarks, and we show this breaks under real deployment conditions. On a 38-class crop disease benchmark split pathologically across ten clients with 5\% labeled data and 10:1 imbalance, a representative baseline (RSCFed) scores an F1 of exactly zero on six classes. The diseases themselves are not the obstacle. The threshold is: it never lets the model learn from them. We propose AgriSSFL, which recalibrates the pseudo-labeling threshold per class from each client's own mean predicted confidence at zero extra communication cost, and aggregates client updates by similarity to the consensus direction rather than by sample count alone, both choices made with low-bandwidth, resource-constrained deployment in mind. AgriSSFL recovers all six abandoned classes, F1 rising from 0.000 to between 0.42 and 0.96, improves macro-F1 by 19.79 points over RSCFed, and improves 33 of 38 classes overall, on a lightweight MobileNetV2 backbone suited to the low-end hardware such deployments would realistically use. Code and reproducibility artifacts are public.