Bridging Learned Visual Perception and Symbolic Belief-Space Planning via Probabilistic Grounding
Abstract
In partially observable settings, agents must act without full knowledge of the world state, and obtaining grounded and verifiable symbolic plans under such uncertainty remains a key challenge. Recent work integrates Vision-Language Models (VLMs) to bridge perception and symbolic reasoning, either mapping images directly to action sequences (VLM-as-planner) or grounding observations into symbolic predicates for off-the-shelf planners (VLM-as-grounder), ignoring uncertainty in the planning process. We introduce a third paradigm, VLM-as-probabilistic-grounder, which captures the uncertainty of VLM predicate groundings as a probability distribution over symbolic states, enabling belief-space planning and robust plans under uncertainty. Experiments in simulated household robot settings show improved robustness and task success over deterministic grounding.