SEISMOS: A Statistical Signal Detection Framework for Semantic Chunking
Abstract
Chunking is a hidden bottleneck in dense retrieval and retrieval-augmented generation: it determines the units that can be indexed, retrieved, and ultimately used as evidence. Yet most systems still segment documents using fixed token windows or globally thresholded semantic similarity drops, treating chunking as an engineering heuristic rather than a statistical decision. We propose SEISMOS, a statistical signal detection framework for semantic chunking. SEISMOS models consecutive sentence-embedding cosine similarities as a document-level signal and derives boundary decisions from the null hypothesis of no semantic transition. Across four development BEIR corpora, we establish three corpus-invariant properties: bimodal document structure, rapid survival decay of shallow local minima, and positive lag-1 autocorrelation. These properties show that semantic boundaries should be detected by variance-normalized deviations, not by document means or global similarity thresholds. Accounting for autocorrelation leads to a normalized discrete Laplacian detector that identifies significant semantic valleys through a single interpretable decision rule. Evaluated on five BEIR benchmarks, SEISMOS consistently improves over fixed-length, recursive, and production semantic chunking baselines under one fixed operating configuration. Without corpus-specific retuning, the same detector transfers to a held-out TREC-COVID corpus. Our results suggest that the boundaries needed for effective retrieval are already encoded in embedding-similarity signals, and that principled, efficient, LLM-free chunking can be obtained by detecting them statistically.