Two Kinds of Nothing: Training Composition Decides Which Absence a Language Model Reads
Samridh Aggarwal ⋅ Leon Janauschek
Abstract
A record saying that X and Y were never tested together reports a search that came up empty. A record that lists the pair but never says whether anyone compared them reports nothing at all. The two are not the same statement, but they answer the same question the same way, so a model that responds differently is reading the text, not the evidence. Which of the two a model gets right depends on its training data, not its size. We train a 14.4M-parameter transformer from scratch on a synthetic corpus, varying one quantity: the fraction of no-finding cases written as an explicit gap rather than left silent. Trained only on stated gaps, models read stated gaps at 87\% and silence at 14\%; trained only on silence, at 55\% and 99\%, with no overlap between the sixteen seeds (exact permutation test, $p = 1.55 \times 10^{-4}$). Neither failure is the price of the other. A silent minority of 3.3\% of the corpus lifts both conditions above 96\% at no cost to stated-gap accuracy, though below that share the outcome turns seed-dependent. Brief continued training on the same mixture also repairs already-broken checkpoints: all eight recover, none loses stated-gap accuracy, and most of the recovery lands inside the first 3\% of the original training budget. The same split appears in thirty released models spanning eight families from 0.5B to 8B parameters, with no relation to parameter count, and in a matched Qwen3-8B pair instruction tuning alone reverses it, from 85\% and 31\% to 24\% and 55\%. Post-training does not remove a model's ability to say it does not know. It changes which absence triggers it.
Chat is not available.
Successful Page Load