A Randomized Framework for Validating Benchmark Contamination Claims
Abstract
Benchmark questions can appear in a model's training data without necessarily affecting its evaluation performance. A model may see an item during training and and not retain it. It may retain the item and still answer by reasoning, leaving its evaluation score unaffected. Or it may retain the item and use this information, overstating the model performance. These are different claims, and evidence for one does not establish another. Telling them apart requires comparing against the same model trained without that item. For a released model, that comparison is not available. We introduce TRACE (Target-Randomized Assessment of Contamination Effects), a method for estimating the effect of benchmark exposure directly. We train two copies of a model under the same recipe and token budget, assigning each benchmark item to one copy or the other at random, so every item is seen by one copy and withheld from its sibling. Comparing the two copies on the same item gives a controlled contrast between exposure and no exposure. From that contrast we measure three effects separately: changes in preference for the exposed solution, reproduction of that solution, and accuracy. We evaluate TRACE across four pretrained model families. Benchmark exposure consistently increases preference for and reproduction of the exposed solutions, but its effect on accuracy varies substantially, from large improvements to no measurable change. Evidence that a model has retained a benchmark item therefore does not by itself show that contamination inflated its score. Where exposure is known, TRACE establishes what contamination evidence is entitled to mean.