Who Watches the Watchers? Semantically-Constrained Reinforcement Learning for Red-Teaming Provenance Intrusion Detectors
Abstract
Reconstruction-based provenance intrusion detection systems (PIDS) detect malicious activity by classifying nodes in system-call provenance graphs from per-edge reconstruction errors produced by benign-trained encoder--decoders. Because these detectors run on the endpoints they monitor, adversarial robustness is a critical deployment concern, yet it remains rarely evaluated systematically. We introduce ProvRL, a reinforcement-learning framework for automated red-teaming of these detectors, and use it to characterise the robustness of four leading PIDS spanning the four encoder families used in the field. ProvRL casts red-teaming as a semantically constrained graph-editing MDP, with a factored autoregressive policy masked by transition relations mined from benign data and a GRU belief state, to model the multi-step consequences of message passing on GNN-based detectors. ProvRL causes targeted false negatives in white-, grey-, and black-box settings on three of four encoder families (GNN, linear, VAE) using 5-99x fewer queries than exhaustive search; direct-reconstruction encoders resist insertion attacks structurally, suggesting an architectural direction for robust detection. No threshold-adaptive aggregation policy maintains both a deployable false-positive rate and meaningful evasion resistance, leaving adversarial retraining as the only viable defence on the vulnerable families. ProvRL trains on CPU within hours and produces attack chains that replay as real Linux syscalls. We release ProvRL as an open-source red-teaming tool for PIDS.