Give the Model a Finger: Developmentally Aligned Evaluation of Counting in Vision-Language Models
Abstract
Vision-language models (VLMs) are exact on small sets and drift into estimation on larger ones, a profile reminiscent of children before they master counting. Children cross the barrier with epistemic actions: pointing at each item while reciting number words, moving counted items aside, partitioning the array. We give VLMs the same actions as dumb image tools with no visual intelligence of their own (a numbered tag drawn at a coordinate, a zoom, an eraser, a grid) in a training-free agent loop. Our central claim is about measurement, not intervention: a model that counts by marking leaves a normally hidden serial procedure on the image, and the developmental principles of counting (used as measurement theory, not as a claim that VLMs are children) turn those marks into an interpretable audit of serial visual reasoning. This developmentally aligned evaluation operationalizes the classic analyses of children’s counting: point-level ground truth classifies every tag as on-target, duplicate, phantom, or omitted (one-to-one correspondence); the answer is checked against the number of tags (cardinality); and order irrelevance is tested by forcing opposite traversal orders. Across 9 models, 682 images (a controlled dot-array battery; FSC-147, 8–80 objects; PixMo-Count, 2–10), and 7 harness conditions we find that (i) the exact-counting range is a model-specific signature spanning nearly two orders of magnitude (R_80 from 3 to 150 dots, the latter located only by extending the battery to 200); (ii) as an intervention the finger mostly fails: tags help five of nine models on small-count photographs (CI-solid for one), hurt imprecise pointers, and rarely beat a verbal count-region-by-region routine on dense photographs; (iii) as measurement it succeeds: pointing precision tracks which models benefit, dense scenes fail by omission (consistent with children’s keeping-track errors), and models rarely revise after seeing their own tags; and (iv) the principles dissociate: models follow an instructed traversal order yet often change their answer when it reverses, beyond a same-order repeat control, and children’s regular-over-scattered advantage appears only under the routine, not under estimation. We release the harness, the stimuli, and all trajectories.