Symbolic Command Grounding over Frozen Neural Perception: When Does a Rule Layer Help Referring Segmentation?
Abstract
Embodied agents receive commands such as segment all red capsules left of the blue triangle'' ordo not segment the pink stars'' and must ground attributes, anchors, relations, multiplicity and action polarity in perception. We study a minimal neuro-symbolic interface: a deterministic parser and a text-only scope guard elicit a symbolic command structure, a frozen referring-segmentation model (CLIPSeg or Grounded SAM) supplies score maps for the command, the target phrase and the anchor phrase, and explicit geometric rules over connected components select the output, abstaining to the native mask whenever a symbolic decision is unsupported. We validate the interface with paired correctness - two commands over one image that differ in exactly one requirement must both be obeyed - on a 1,200-scene procedural suite with exact instance masks, matched flat/rich renders and five seeds, and on all 49,492 gRefCOCO expressions. The rule layer raises paired correctness by 5.7-16.8 points and survives texture, rotation and occlusion, but the size and even the direction of the nuisance shift depend on the neural interface; anchor geometry carries the whole gain, and an opposite-relation counterfactual consistency check adds cost without benefit. On natural language the symbolic grammar reaches only 3.15% of expressions, widening it can lower utility while doubling harms, and 93-95% of harmful interventions arise where the symbolic union rule meets neural connected components rather than instances.