Little Models, Big Literature: Rebuilding a Review Table with Small Language Models on a Laptop
Abstract
Agentic literature review repeats one narrow extraction task across every document, so per-document cost decides feasibility. The documents may also be proprietary, making local execution a privacy requirement as well as an economic one. We ask whether small open-weight models on a laptop can rebuild a published review table from a 1346-document corpus, and how to compare them fairly against frontier models that already hold the answer in memory. The most contaminated model we probe reproduces 34 of the 39 recoverable values closed-book, with no documents provided. Our protocol therefore controls three axes: contamination, input budget, and cell recoverability. Under matched inputs, more context is not generally better, recoverability tiers derived from a frontier-heavy committee do not transfer to small models, and a provenance audit exposes right-number, wrong-quantity errors invisible to recall. Under these controls, a 20B model running both stages on a laptop reconstructs 22 of the 39 recoverable cells, with no frontier model in the chain and no API bill.