GroundBench: A Factorized, Counterfactual Benchmark for Locating VLM Affordance Failures
Abstract
A companion evaluation found that naming a target part in a manipulation prompt raises action accuracy by 0.32–0.63 across eight vision-language models (VLMs). That intervention, however, supplies a variable a real system must infer. We in- troduce GroundBench to distinguish target identity, target region, and mechanical information through six branch-and-merge conditions and a within-object coun- terfactual re-ask. Across three OpenAI models and 1,068 predictions, supplying the target region without its identity leaves action accuracy at or below the 0.53 majority baseline (0.26, 0.26, 0.53), although models largely reproduce the sup- plied region. Supplying identity without location yields 0.74, 0.68, and 0.68. Ev- ery above-baseline gain in this curated set occurs when the supplied target-part category itself determines the action. A no-vision control costs GPT-5 nothing on these conditions and improves two of three scores, evidence consistent with category-to-action association contributing substantially to the companion result; GPT-4o mini drops on one condition, so the interpretation is not universal. Adding joint type and motion axis never improves accuracy across six model–stratum comparisons. On a frozen action-contrast set of 32 objects and 74 informative pairs, GPT-5 attains 0.86 pair-weighted compliance but fails all nine observed push→lift-vertical cases; lift-vertical is itself rarely emitted. Ground- Bench is a diagnostic prerequisite for targeted post-training: it identifies which supplied information changes affordance behavior and when apparently grounded performance can be reproduced through textual shortcuts.