Reference, Retrieve, Resolve: How Small a Language Model Can Read a Reference in Context
Abstract
On-device assistants mostly do one shape of task: a reference in context, a short question, a lookup and a rule. We build a narrow eval of that shape and ask how small a model can do it. The reference is the Generation I Pokémon type chart and the typing of all 151 species; the answer is the damage multiplier, a six-way choice with a code-generated key. Nine open models from 1B to 235B and one closed ceiling model run a pinned 400-item set at three epochs, one host per model, with a pre-registered second draw of 400 and two controls. Four findings. On this ladder no model at 8B or below that answers in one pass reads the chart as a grid; the cleanest step is within one family, Gemma 3 from 0.28 at 4B to 0.51 at 12B. Rendering the chart as one line per attacking type raises nine of ten models, by up to 0.18, up to a rung. A reasoning trace is worth more than any rung: a thinking 8B (0.850) beats every one-pass model on the ladder including a 235B mixture, and parameter counts did not place the two mixtures. And the score is reading, not memory: recall alone is 0.14 to 0.28 through 27B, and relabeling the fifteen types, which makes memory useless, moves no model by more than 0.06 at a matched token budget. Two swings that looked like capability were harness artefacts. Harness, key, item sets, registration files and ledger are released.