Towards Unconstrained Pedestrian Attribute Recognition: Evaluating Domain Shift and Open-World Constraints
Dattaraj Salunkhe ⋅ Vaibhav Rathore ⋅ Biplab Banerjee
Abstract
Pedestrian Attribute Recognition (PAR) is normally trained and tested on the same attribute vocabulary under the same visual conditions. Deployment gives neither: cameras go up in scenes the model never saw, and operators ask for attributes nobody annotated. We study both problems jointly and formalize the setting as Generalized Zero-Shot Pedestrian Attribute Recognition under Domain Shift (GZS-DG-PAR). Attribute names are known in advance, but visual labels for a subset of them are withheld during training, and evaluation happens on scenarios held out from training. We build a benchmark by partitioning the large-scale MSP60K dataset along its native scenarios, avoiding the artifacts of merging separate PAR corpora. Our analysis surfaces an invariance-flexibility trade-off: vision-language models can name unseen attributes but their text-image alignment degrades under covariate shift, while domain-generalization methods stay robust yet remain tied to a fixed classifier head. We propose PerceptPAR, combining attribute-guided sparse cross-attention, a hybrid semantic-visual graph, and an evidential prediction head. PerceptPAR attains the highest Overall mean Accuracy (0.6938) using roughly $3\times$ fewer total parameters than the strongest baseline. Most of this gain comes from seen attributes; zero-shot ranking on unseen attributes stays close to chance for every method we evaluate, and we characterize why.
Chat is not available.
Successful Page Load