Physical World Model
Abstract
Modern computer vision and embodied AI systems must now \emph{act} in the physical world, not merely describe it. Yet most pipelines still interpret it through RGB pixels, ignoring the structure that governs how objects move, sound, deform, and respond to contact. We argue that physical world understanding requires AI systems to jointly reason about three coupled pillars, including: \textbf{(1) Physical Geometry:} 3D/4D structure, articulated and deformable motion, pose, and spatial scene understanding. \textbf{(2) Physical Characteristics:} material properties, dynamics, deformability, affordances, and physical interactions. \textbf{(3) Physical Sensors:} non-RGB modalities including tactile, force/torque, proprioceptive, audio, RF, depth, inertial, and event-based sensing. The \textbf{1st Workshop on Physical World Understanding (PhysWorld)} brings together researchers from computer vision, robotics, graphics, multimodal learning, haptics, audio, and physics-based simulation to address a central question: \textit{How can AI systems perceive, represent, and reason about the physical world beyond appearance?} \textbf{Importance and timeliness.} Recent advances in foundation models, embodied AI, world models, multimodal sensing, and differentiable simulation have matured the methods needed to treat physical perception as a unified problem spanning geometry, dynamics, materials, and sensing. At the same time, applications in robotics, AR/VR, autonomous systems, digital twins, scientific imaging, and human-AI interaction increasingly require models that understand the physical consequences of actions rather than merely visual appearance. Despite rapid progress, these research directions remain fragmented across separate communities and venues. Without a coordinated forum now, robotics, vision, and audio sub-communities will independently entrench incompatible benchmarks, ontologies, and evaluation protocols for physical-world foundation models, locking in fragmentation that retroactive standardization rarely undoes. PhysWorld convenes these communities to look beyond pixels, establishing a common forum before that fragmentation hardens. \textbf{Expected outcomes.} Shared evaluation protocols across the three pillars; a community-validated articulation of physical foundation models that integrate vision with touch, sound, proprioception, and dynamics; a cross-disciplinary author cohort publishing together across vision, robotics, graphics, and multimodal sensing; and an accessible entry point for early-career researchers into multimodal physical AI. \noindent\textbf{Topics of interest include, but are not limited to:} \textbf{(a) Physical Geometry:} 3D/4D reconstruction, articulated and deformable scene understanding, geometry-aware world models, and physically grounded view synthesis. \textbf{(b) Physical Characteristics:} estimating material and physical properties from vision and multimodal sensors (mass, friction, stiffness, elasticity, deformability, affordances); physics-informed learning, differentiable simulation, contact-rich interaction modeling, and generative models of physical dynamics. \textbf{(c) Physical Sensors:} multimodal sensing and sensor fusion using tactile, force/torque, proprioceptive, RF, audio, depth, IMU, and event-based signals. \textbf{(d) Cross-cutting:} embodied world models, robot manipulation, sim-to-real transfer, multimodal simulators, and physically grounded reasoning for autonomous agents; benchmarks, datasets, evaluation protocols, and responsible deployment of multimodal physical AI systems.