Uncertainty-Aware Listwise Reinforcement Fine-Tuning for Fine-Grained Visual Classification
Abstract
Fine-grained visual classification (FGVC) with Multimodal Large Language Models (MLLMs) remains challenging, as visually similar subordinate categories inherently induce classification uncertainty due to subtle inter-class differences. Existing reinforcement learning methods typically require the model to generate a single class name and optimize it with sparse correctness rewards. However, a seemingly incorrect prediction can still provide highly informative signals if it falls within the same local confusion set as the ground-truth class. By treating all non-exact matches as equally wrong, the rigid single-answer paradigm overlooks these valuable near-misses. Consequently, it fails to capture the model's intrinsic uncertainty and provides limited learning signals. To address these limitations, we propose List-GRPO, an uncertainty-aware listwise Group Relative Policy Optimization framework. Rather than enforcing a single prediction, our method prompts the model to adaptively generate a ranked candidate list, explicitly capturing its predictive uncertainty. Furthermore, to mitigate the advantage vanishing when online exploration completely misses the target, we introduce a dynamic ground-truth anchored response strategy. We synthesize an off-policy anchor by prepending the ground truth to the most frequent hard negatives mined from online samples. To prevent the anchored response from dominating learning, we isolate the advantage estimation of online samples from the anchored response and use a gating coefficient to modulate its contribution. Extensive experiments on multiple fine-grained benchmarks demonstrate that List-GRPO significantly improves Top-1 accuracy while maintaining high ground-truth coverage with only about three candidate predictions on average.