SeismoDet: Modeling Local Perturbation Propagation for Training-Free LLM Detection
Abstract
Large language models (LLMs) can generate increasingly fluent text that is difficult to distinguish from human writing, creating a need for reliable detection methods that do not require labeled training data. Existing training-free approaches exploit intrinsic properties of token probabilities, entropy, likelihood ranks, or their temporal and spectral characteristics. However, these methods largely focus on the magnitude, distribution, or frequency content of token-level fluctuations, while paying less attention to how a local change evolves and propagates through the subsequent sequence. This distinction is important because human- and machine-generated text can exhibit substantial overlap in their marginal probability statistics, making static or fluctuation-based signals insufficient in challenging settings. We introduce SeismoDet, a training-free detector that addresses this gap by modeling token log-probability trajectories as propagating signals through a seismic-wave-inspired framework. In seismology, the observed response to a disturbance is determined not only by the initial event but also by how the resulting energy propagates, scatters, and attenuates through the underlying medium. The subsequent scattered waveform, or coda, carries information about the structure and heterogeneity of that medium. We translate this principle to language generation by treating a local change in the token likelihood trajectory as a perturbation and examining how its signal evolves across subsequent tokens. SeismoDet therefore captures transition propagation, scattering behavior, and coda stability, providing information that is not directly represented by conventional likelihood or temporal statistics. Our central hypothesis is that these propagation dynamics reflect differences in the underlying generation process. Human-written text tends to exhibit more heterogeneous and irregular transitions, leading to stronger local disturbances and greater variability in their subsequent trajectories. In contrast, LLM-generated text often exhibits smoother transition dynamics and more stable post-transition behavior. Thus, rather than considering only whether individual tokens are likely or unlikely, SeismoDet asks how a probability perturbation propagates through the sequence and what structure remains after the transition. This provides a complementary perspective to existing temporal- and frequency-domain approaches. We evaluate SeismoDet across diverse datasets, source models, decoding strategies, and both black-box and white-box detection settings. In the white-box setting, SeismoDet achieves an average AUROC of 95.25%, improving over SpecDetect by 1.25 percentage points and Lastde by 3.41 points. In closed-source black-box detection, it achieves 83.80% average AUROC, outperforming Lastde by 1.28 points. On the HC3, RAID, and M4 benchmarks, SeismoDet achieves 88.95%, 87.84%, and 89.16% AUROC, respectively. These results demonstrate that modeling propagation and coda behavior provides a complementary and effective signal for training-free LLM-generated text detection. More broadly, this work shows that physics-inspired modeling can provide useful inductive biases for understanding language-generation dynamics. Future work could explore other physical principles to build a more general framework for robust, training-free LLM-generated text detection.