PhysFormer: Learning to Simulate Mechanics in World Space
Abstract
We present PhysFormer, a physics-grounded diffusion transformer for generating 4D multi-object mesh dynamics directly in world coordinates. Rather than predicting future frames in pixel space or rolling out next-step system states autoregressively, PhysFormer models physical evolution as full-trajectory coordinate diffusion. Given initial per-vertex positions, velocities, and material conditions, it generates future mesh vertex trajectories in a single denoising process, producing physically plausible object-object and object-environment interactions without hard-coded constraints, simulator priors, or learned shape latents. PhysFormer adopts a DiT-style backbone with factorized temporal, spatial, and object-level attention to capture coherent structure across time, vertices, and objects. Trained on over 100k collision-rich simulated trajectories, it models rigid and deformable multi-object dynamics, generalizes to unseen real-world geometries and larger object counts, and substantially outperforms autoregressive baselines in trajectory accuracy, rigidity preservation, and momentum-based physical consistency. Our results position coordinate-space diffusion as a promising step toward view-invariant, geometry-level world models for robotics, graphics, and physical design.