Evaluating The Reliability of Feature-Attribution for Adversarial Evaluation of Phishing Detectors
Abstract
We evaluate whether feature-attribution-based adversarial LLM rewriting is a reliable evaluation methodology for finding stylistic vulnerabilities in black-box phishing detectors. We test this for two large language model (LLM) and bidirectional encoder representations from transformers (BERT) phishing detectors using generic and feature-targeted rewriting. In this paper, we generate a generic rewriting baseline and train XGBoost surrogate models to predict evasion from 50 features and their deltas. We then use SHAP to analyze feature attributions and generate targeted datasets based on the most influential features for each significant surrogate. Our findings suggest that three of four surrogate models learn significant feature-evasion relationships that partially generalize across newly-generated datasets. Targeted rewriting generally increases evasion relative to the generic baseline, but the increase is not detector-specific. Additionally, LLM rewriting frequently fails to move the targeted features in the intended direction. This suggests that feature-based surrogates can be useful for predicting black-box detectors' evasions, but converting SHAP feature attributions to LLM rewriting instructions alone is insufficient to reliably generate targeted adversarial datasets. An evaluation on held-out emails supports the cross-detector transferability finding but not the consistent match between predicted and actual evasion rank. This suggests that our investigated SHAP-attribution-based LLM rewriting methodology is not reliable for detector-specific evaluation.