Seeing Together, Acting Apart: Shared Environmental Understanding for Multi-Robot Navigation
Abstract
Multi-robot embodied systems are commonly expected to collaborate by coordinating their behaviors: agents may communicate, share experience, divide tasks, avoid conflicts, or jointly decide what to do next. This paper asks a complementary question: can collaboration emerge before decision making, by first building a shared understanding of the environment? We introduce Shared Environmental Understanding, a policy-agnostic representation layer that aggregates partial egocentric observations from multiple robots into a common latent world representation. This formulation shifts the starting point of collaboration from agent-centric traces to environment-centric representation, allowing different robots and downstream policies to benefit from shared world knowledge without being forced into a centralized controller. We instantiate this idea with SEER, a shared environment encoding and retrieval framework that learns to turn multi-robot, multi-view observations into reusable navigation context. SEER is designed not as a new navigation policy, but as an upstream environmental understanding module that can be plugged into heterogeneous agents, including both specialized VLN models and general vision-language model agents. We validate SEER through a comprehensive set of representation-level and task-level studies, including latent inverse dynamics analysis, standard VLN evaluation, controlled paired multi-robot navigation, comparisons against action-, history-, and policy-sharing alternatives, and real-world robot deployments. Across these settings, SEER preserves single-agent navigation ability while improving individual and collaborative success in multi-robot scenarios. These findings support a simple but underexplored hypothesis: for embodied cooperation, robots may benefit from sharing the environmental understanding before sharing the policy.