OCC4M: Grounded Object-Centric Memory for Neuro-Symbolic Manipulation
Abstract
Embodied agents must connect learned perception to persistent identities, relations, and actionable locations even when relevant objects are no longer visible. We present OCC4M ("Occam"), an object-centric 4D memory that serves as an explicit interface between neural perception, language-based task reasoning, and learned manipulation skills. Deterministic geometric association maintains world-frame tracks and temporal, motion, and containment relations; a vision-language model (VLM) selects identities and targets from this state, while a frozen, history-free executor acts on grounded visual cues. Across seven simulation conditions and 350 episodes, OCC4M achieves 96.6\% memory success and 88.9\% end-to-end success, versus 54.6\% and 57.7\% for a raw-history VLM given the complete observation history and the same executor. A controlled viewpoint change yields 100\% memory success for OCC4M and near-zero success for the baseline. On 20 real-world Franka episodes, OCC4M achieves 85\% joint memory accuracy versus at most 30\% for the baseline, and completes 45\% of full two-stage physical tasks. A perturbation demonstration shows identity-bound target updates without a new VLM call. These results support persistent grounded state as a practical interface linking task reasoning to action in embodied neuro-symbolic systems.