StemBind: When MLLMs Get Lost Between Rules and Instances in Abstract Visual Reasoning
Abstract
Multimodal large language models (MLLMs) often know the rule but pick the wrong answer: on abstract visual reasoning (AVR) tasks, a model can correctly describe what it sees and correctly name the underlying pattern, yet still fail to choose the matching candidate. Existing AVR benchmarks cannot detect this gap because they collapse perception, rule induction, and answer selection into a single right-or-wrong signal. We introduce StemBind, a shared-stem diagnostic benchmark that probes the same visual stem with three aligned questions: Perception (what is in the image), Rule (what pattern governs it), and Full (which option completes it), so that a final-answer error can be attributed to a specific sub-step on the same evidence. StemBind contains 2,298 curated knowledge-light stems across nine auditable visual operations, totaling 19,533 P/R/F tasks, with each full item annotated by Sternberg's four reasoning stages: S1 Encode, S2 Infer, S3 Map, and S4 Apply. Evaluating 24 frontier MLLM configurations (proprietary and open-source) yields four findings. (i) The R-F chasm. Rule accuracy exceeds full-item accuracy on 22 of 24 models, so most failures happen after the rule has been identified. (ii) A persistent binding gap. Even when P and R are both correct on the same stem, models still answer F incorrectly 51.2% of the time. (iii) The bottleneck is S3. Process diagnostics and Stage-wise Stimulus Augmentation (SSA) localize the dominant failure to rule-to-instance mapping, the step that binds an inferred rule to the right candidate. (iv) Scaling and thinking do not help. Neither scaling up model size nor enabling explicit thinking mode reliably closes the gap; in paired comparisons, thinking mode lifts perception but lowers both rule and full-item accuracy. StemBind reframes AVR evaluation from final-answer ranking to locating where abstract visual reasoning breaks down, and identifies rule-to-instance binding as a concrete next target for vision-grounded reasoning.