Text Before Tokens: A Linguistically-Grounded, Model-Free View of Entropy in Foundation Models
Lewis Mitchell
Abstract
Entropy measures reported for foundation-model outputs -- perplexity, per-token surprisal, logprob-based Shannon entropy -- are computed after text has been segmented by a model-specific byte-pair-encoding (BPE) vocabulary, so the same string can receive different entropy purely as an artefact of which tokenizer sits underneath the model being probed. The Kontoyiannis entropy rate estimator $H_K$ avoids this by operating directly on raw text, with no vocabulary or model access required, and this paper argues for treating that property as a virtue rather than an incidental feature. We situate $H_K$ within a compression-based tradition in linguistics for measuring language complexity that predates large language models and has not, to our knowledge, previously been connected to model collapse. We then revisit two existing empirical results as case studies: $H_K$-based filtering re- sists fine-tuning collapse where logprob-based filtering does not, and $H_K$ serves as a common observable state for a dynamical-systems account of collapse fit across model scales. Finally, we report a new test retokenizing the same fixed documents under five schemes spanning three distinct segmentation algorithms -- two byte-pair-encoding (BPE) vocabularies, a WordPiece vocabulary, a Uni- gram/SentencePiece vocabulary, and a non-subword whitespace baseline: contrary to our pre-registered expectation, all four subword tokenizers agree closely on the resulting diversity-collapse trajectory, and the whitespace baseline follows the same proportional decline. The paper’s case nonetheless stands, since $H_K$’s independence from tokenizer choice is a guarantee of its construction, not a fact contingent on which tokenizers happen to agree empirically on a given corpus.
Chat is not available.
Successful Page Load