Concept Bottleneck Models in the Wild: Assessing Interpretability under Image Corruptions
Kexin Zhang ⋅ Katharina Prasse ⋅ Mishal Fatima ⋅ Margret Keuper
Abstract
Concept Bottleneck Models (CBMs) route every prediction through a layer of named, human-auditable concepts, which makes their decisions inspectable. To date, CBMs have only been evaluated in the lab setting, however, when employed in the wild, new challenges may arise: We ask what happens to a deployed CBM when common image corruptions of real sensors alter its inputs. We analyse three CBMs and find that image corruptions reduce model accuracy by $3-67$%, depending on corruption type and severity. Two findings organize the paper. First, image corruption reorders concept activations. While the CLIP embedding moves by only a bounded amount even at the harshest severity (cosine $0.59$ or above), the standardized activations the classifier reads are re-ranked so thoroughly that the single top concept changes in $70$ to $100$% of images. Second, the reordering is uneven across concepts, so some concepts are more stable than others. We propose a per-concept stability score $d_k$. Our experiments show that pruning instable concepts improves accuracy on corrupted images, while having little impact on clean accuracy. Code will be released upon acceptance.
Chat is not available.
Successful Page Load