Emergent retokenization symmetry in large language models: phenomenology and applications
Kanishk Jain ⋅ Matthew Day ⋅ Tankut Can
Abstract
Tokenization introduces representational redundancy: under a fixed vocabulary, every byte string admits many valid token encodings, or segmentations, that decode to the same surface string. Most language model tokenizers break this symmetry by returning a canonical segmentation, and training only on canonical segmentations gives little reason to expect downstream symmetry. Nevertheless, we find that this symmetry partially emerges during training. We probe it through experiments on token compositional understanding, representation diversity, and task-focused benchmark performance. Our main intervention is $\textbf{retokenization}$ - replacing a prompt's canonical tokenization with another valid segmentation while preserving its bytes exactly. Unlike other prompt perturbations, retokenization isolates segmentation effects without changing syntax, semantics, or surface form. We use it to study sensitivity and robustness to semantically identical inputs across pretraining and post-training, and use this emergent partial symmetry as an inference-time sampling axis. While temperature sampling generates diverse outputs from the model using its next-token probability distribution, retokenization generates diversity from the model's internal computations through semantically equivalent input representations. We also find that while retokenization can hurt performance on easy problems, it can also recover solutions that conventional sampling sometimes misses. Overall, retokenization helps probe compositional understanding and prompt sensitivity, and can be used as a novel sampling strategy.
Chat is not available.
Successful Page Load