Transferring Visual Explainability from Self-Explaining to Prediction-Only Vision Transformers via Task Arithmetic
Abstract
In image classification scenarios where both prediction and explanation efficiency are required, self-explaining models that perform both tasks in a single inference are effective. However, for users who already have prediction-only models, training a new self-explaining model from scratch imposes significant costs in terms of both labeling and computation. This study proposes a method to transfer the visual explanation capability of self-explaining Vision Transformer (ViT) models learned in a source domain to prediction-only VLM-based ViT models in a target domain via task arithmetic, without target-side explanation supervision. The proposed method endows explanation capability by adding an \emph{explainability vector} induced by explanation supervision in the source domain based on task arithmetic framework. Experiments on ten diverse target datasets show that a single explainability vector learned on ImageNet-1k augmented with patch-level explanation supervision transfers consistently, improving explanation quality while largely preserving classification accuracy. Beyond the transfer itself, we further investigate whether transfer success can be anticipated \emph{a priori} from model-internal features alone (without target-side explanation supervision), and find that a small set of such features carries non-trivial signal about the transfer outcome.