Concept Bottleneck Models in the Wild: Assessing Interpretability under Image Corruptions
Kexin Zhang ⋅ Katharina Prasse ⋅ Mishal Fatima ⋅ Margret Keuper
Abstract
Concept Bottleneck Models (CBMs) route every prediction through a layer of named, human-auditable concepts, which makes their decisions inspectable. To date, CBMs have only been evaluated in the lab setting, however, when employed in the wild, new challenges may arise: We ask what happens to a deployed CBM when common image corruptions of real sensors alter its inputs. We analyse three CBMs and find that image corruptions reduce model accuracy by $3-67$\%, depending on corruption type and severity. Two findings organize the paper. First, image corruption \emph{reorders} concept activations. While the CLIP embedding moves by only a bounded amount even at the harshest severity (cosine $0.59$ or above), the standardized activations the classifier reads are re-ranked so thoroughly that the single top concept changes in $70$ to $100\%$ of images. Second, the reordering is uneven across concepts, so some concepts are more stable than others. A per-concept pre-softmax stability score $d_k$ measures how far each concept moves under corruption, and keeping only the most stable concepts, at a fixed bank size matched against a random subset of the same size, recovers $1.5$ to $3.3$ points of corrupted accuracy across the three models while keeping clean accuracy constant or even improving it. Our analysis of $d_k$ shows it to be part genuine shift-robustness and part crop-acquisition geometry for image-based banks, so the robust subset is best used as a diagnostic and sparsity instrument. Code will be released upon acceptance.
Chat is not available.
Successful Page Load