Explanation-first Explainable AI
Abstract
Current methods for explaining AI models are often incomplete or unfaithful to the mechanisms that produced a decision, and frequently fail to convey the insights an audience needs to understand how the decision was reached. We argue this should be addressed in the design objective. Under post-hoc XAI, models are trained solely for performance and explainability comes after; with interpretable models, we balance both performance and understanding. We propose explanation-first XAI, which inverts these priorities: rather than fitting explanations to the model, we design solely for the explanation we would like, and then build a model that implements it. To make explanation quality measurable, we evaluate the explanation as a communication tool: does a second party understand it? and introduce a mental-model framework for evaluating understanding and improving explanations. We demonstrate the viability of this approach on ARC-AGI. This work offers an alternative to post-hoc XAI and interpretable models, by directly prioritizing the quality of communication provided by the explanation.