IDEA: Unwrapping Visual Black-box Models by Interaction Decomposition
Abstract
Concept-based explanation methods have emerged as a prominent framework for interpreting deep neural networks using human-understandable concepts. However, existing approaches are limited to assigning a scalar importance score to a concept, failing to capture the interaction between the concept and its visual context. We argue that overlooking this interaction leaves a predictor's reasoning only partially understood. To address this gap, we introduce IDEA, a novel global post-hoc explanation framework that explicitly models the interaction between a concept and its visual context. Through this modeling, IDEA decomposes concept-level predictive information into three atomic components: uniqueness (independent concept contribution), redundancy (information shared with the context), and synergy (information arising strictly from concept-context interaction). Consequently, IDEA extends interpretability beyond merely identifying influential concepts to diagnosing how a predictor reasons -- for instance, flagging potential context dependencies that scalar scores cannot reveal. IDEA requires no task-specific concept annotations and no access to classifier internals. Evaluations on controlled and real-world settings demonstrate that IDEA can surface reasoning patterns invisible to traditional scalar concept scores.