Do Deepfake Heatmaps Identify Image Changes? Evaluating B-cos and Post-hoc Explanations
Tilo Schütz ⋅ Toby Fuchs ⋅ Kai Reffert ⋅ Linus Neumann ⋅ Nils Becker ⋅ Robin-Kiara Braun ⋅ Margret Keuper ⋅ Steffen Jung
Abstract
Deepfake detectors can support information integrity, but their scores do not reveal whether they use manipulation evidence or shortcuts. We study B-cos networks, whose predictions directly yield heatmaps, instead of explaining conventional detectors afterward. We train four detector families on FaceForensics++ (FF++) and evaluate them on its test set, and 18 out-of-distribution evaluation sets spanning different data sources and manipulation methods. We additionally assess their explanation maps using three pointing-game protocols. Our new Pixel Pointing Game uses aligned real--fake pairs to retain half of the annotated manipulation with the largest pixel differences. B-cos Xception matches standard Xception on FF++ while improving mean accuracy on five established benchmarks. When tasked with detecting the only fake among 9 images, its built-in heatmap places $.95$ of its positive evidence in the manipulated image, versus $.41/.49$ for Grad-CAM/++ on the same detector. While Grad-CAM/++ finds explanations that score almost perfectly inside broad face masks, its strongest explanations are outside the manipulated area. Across 15 B-cos settings, the built-in explanation performs best in 14 cases after accounting for the smaller target. While heatmaps provide useful audit information, they must be accompanied by performance tests on new data to ensure generalization.
Chat is not available.
Successful Page Load