Multi-bit LLM Watermarking with Certified Semantic Distortion
Abstract
Watermarking has emerged as a fundamental mechanism for tracing the provenance of Large Language Model (LLM) outputs. While multi-bit watermarks are highly desirable for embedding rich metadata, they suffer from an inherent trade-off between statistical detectability and text quality degradation. Crucially, existing multi-bit schemes lack theoretical quality guarantees, leaving them susceptible to unbounded semantic distortion. To address this, we introduce CSD, the first multi-bit watermarking framework with certified semantic distortion bounds. In CSD, we derive a closed-form upper bound on the Kullback-Leibler (KL) divergence, mathematically guaranteeing strict limits on semantic distortion at every generation step. Furthermore, rather than relying on static payloads, CSD employs a dynamic payload to maximize statistical detectability. Extensive empirical evaluations across multiple datasets demonstrate that CSD achieves robust detectability while maintaining provably safe text generation.