Stop Using Plausibility as the Criterion for Explainable AI
Abstract
Explainable artificial intelligence (XAI) is motivated by the problem of making AI predictions understandable, transparent, and responsible, as AI becomes increasingly impactful in society and high-stakes domains. The evaluation and optimization criteria of XAI are gatekeepers for XAI algorithms to achieve their expected goals and should withstand rigorous inspection. To improve the scientific rigor of XAI, we conduct the first comprehensive and critical examination of a common XAI criterion: plausibility. Plausibility assesses how convincing the AI explanation is to humans, and is usually quantified by metrics of feature localization or feature correlation. Our examination shows that plausibility is invalid to measure explainability, and human explanations are not the ground truth for XAI, because doing so ignores the necessary assumptions underpinning an explanation. One fundamental assumption is that AI explanation is supposed to behave adversarially, analogous to auditors' or doctors' roles of identifying potential problems. Evaluating or optimizing XAI to generate more plausible explanations is analogous to incentivizing (i.e., bribing) auditors/doctors to generate more positive reports, which is different from encouraging AI models to learn more plausible features (analogous to encouraging companies/patients to improve their conduct/health). Our examination further reveals the consequences of using plausibility as the XAI criterion, including increasing misleading explanations that manipulate users, deteriorating users' trust in the AI system, undermining human autonomy, being unable to achieve complementary human-AI task performance, and abandoning other possible approaches of enhancing understandability. Due to the invalidity of measurements and the unethical issues, this position paper argues that the AI community should stop using plausibility as the criterion to evaluate and optimize XAI algorithms. We also delineate new research approaches to improve XAI in trustworthiness, understandability, and utility to users including complementary human-AI task performance.