[short] Observable Symbolic Validation and Capacity-Aware Recovery for Safer Vision-Language Robotic Sorting
Abstract
Industrial robotic sorting requires a robot to recognize incoming components, assign them to the correct destination, and execute the corresponding placement safely. Vision-language models (VLMs) provide a flexible interface for such decisions, but incorrect output can propagate directly to unsafe physical actions. We study a lightweight neuro-symbolic execution layer that contains such errors without modifying the VLM's original decisions. The system validates high-level sorting commands using observable appearance and scene constraints and redirects rejected items to capacity-limited recovery locations. We evaluate the approach in a simulated seven-class sorting cell using three open-weight VLMs under paired execution conditions. With Llama-3.2-11B Vision, validation eliminated all 22 unsafe shelf placements, while recovery increased safe handling from 65.1\% to 98.4\%. Gemma-4-31B and PaliGemma-3B also reached zero unsafe shelf placements; capacity and verifier-noise ablations exposed limits from recovery capacity and false rejection. These results support post-decision symbolic validation and recovery as a modular safety layer for VLM-guided robotic sorting.