Weighing What Matters: Corpus-Derived Token Importance for Language Model Objectives
Abstract
Standard autoregressive language model training optimizes a uniform cross-entropy objective that treats every next-token prediction equally. Since natural language follows a highly skewed word-frequency distribution, this objective allocates a disproportionate share of optimization to frequently occurring, but often semantically weak tokens. The challenge is to identify informative prediction targets and redistribute optimization effort accordingly without compromising the language modeling objective. We introduce a fixed, corpus-derived objective-weighting framework: document-level lexical-specificity scores are projected onto next-token targets before training, computed once from corpus statistics rather than model state, and retain dense supervision. We instantiate it with TF-IDF scores projected onto GPT-2 BPE targets. Across general-domain and biomedical corpora, TF-IDF weighting consistently improves objective-aligned evaluation and downstream performance, increasing PubMedQA accuracy from 39.0\% to 50.7\%. Linguistic analyses localize these gains primarily to highly weighted nouns and named entities. Fine-tuning across four Pythia model sizes (14M-160M parameters) further demonstrates that the benefits of weighting transfer across architectures, tokenizations, and capacities, while the general-domain advantage diminishes with increasing model size. Overall, corpus-derived objective weighting provides a simple mechanism for steering limited optimization capacity toward corpus-informative prediction targets, with the largest benefits in small-capacity and domain-specific settings.