Robust and Hard-to-Remove GNN Watermarking via Topological Invariant Perception
Abstract
Graph Neural Networks (GNNs) represent valuable intellectual property, yet existing watermarking schemes primarily rely on OOD backdoor triggers that are susceptible to model pruning, fine-tuning, and distillation. To tackle this challenge, we present InvGNN-WM, which ties ownership to a model's implicit perception of a graph invariant, enabling trigger-free, black-box verification with negligible task impact. By training a scalar head to predict normalized algebraic connectivity on owner-private carrier graphs, ownership is embedded into the model's core reasoning logic rather than exogenous patterns. We provide guarantees for imperceptibility and robustness, and prove that exact removal is NP-complete under monotone decoders. Empirical evaluations across diverse node and graph classification datasets show that InvGNN-WM maintains clean task accuracy while outperforming trigger- and explanation-based baselines in watermark fidelity. Our method remains robust under unstructured pruning, fine-tuning, and post-training quantization, with clear recovery pathways under knowledge distillation.