SCTI: Self-Calibrated Trident Identification of Black-Box LLM Watermarks
Abstract
Black-box watermarked LLM identification has become an important task for watermark auditing. The core challenge arises from inherent black-box constraints that deny access to logits, detector keys, model parameters, and internal watermark settings. For this reason, existing methods have not conducted in-depth exploration on this task, leaving prominent limitations: Fixed-Reference Calibration and Family-Specific Tests. First, existing works use fixed null reference to serve as a baseline for judging the presence of watermarks in LLM outputs, which might misidentify a watermarked LLM as unwatermarked because the null reference is sensitive to the prompt pair, queried model, and sampling configuration. Second, existing statistical-tests based methods adopt distinct statistical features and judgment criteria for different LLM watermark families, making them less suitable when the deployed watermark algorithm is unknown. To remedy the above two limitations, this work proposes SCTI, a Self-Calibrated Trident Identification framework for black-box watermarked LLM identification. Specifically, to handle the first limitation, SCTI constructs empirical null distributions from the queried responses to estimate the generalizable reference instead of relying on a fixed reference. Meanwhile, to address the second limitation, SCTI is configured with a Trident-View Consistency measuring mechanism, which enables a unified pipeline to avoid designing separate tests for individual watermark family. Experimental results verify that SCTI outperforms representative black-box watermark identification baselines across diverse LLMs and watermark algorithms.