Selecting Against Grammar: An Inverted Ranking of Pretraining Recipes by Downstream QA
Deep Gandhi ⋅ Ali Asaria ⋅ Tony Salomone
Abstract
No stage of the pipeline that selects pretraining data measures grammatical competence: candidate recipes are now ranked with small-scale proxy suites scored on downstream question answering. We put a linguistic question to the most prominent such suite, DataDecide: does ranking its 25 data recipes by grammatical competence reproduce the ranking its downstream-QA aggregate produces? Scoring every released checkpoint on minimal-pair grammaticality, we find that at the model size where the signal is statistically resolvable the two rankings oppose each other, at Spearman $\rho = -0.60$ with a bootstrap interval excluding zero, and the recipe ranked first on QA sits near the bottom on grammar. The cost is carried by which quality classifier a recipe was filtered with and not by how much text the filter removed, since FineWeb-style classifiers cost several points of grammar while a plain DCLM filter at comparable retention costs nothing measurable. Because the classifier-filtered corpora are also the highest scoring on QA, the proxy trades grammatical competence for downstream performance. At the smaller size we test, the correlation is not resolvable, which we report as a non-replication and read as scale dependence, since the extremes of the distribution are preserved. That reading sharpens the concern, because these suites are built for the small-scale regime in which the discrepancy is hardest to see. Data-decision suites need a linguistic evaluation axis; without one, optimizing the QA proxy can select against grammar.
Chat is not available.
Successful Page Load