WECO: Equivariant, Generalizable Robot Actions from Object-Centric World Changes
Maggie Meiqi Shao
Abstract
Generalizing manipulation beyond the training distribution remains difficult because robot demonstrations sparsely cover object appearances, geometries, and workspace configurations. We introduce \textbf{WECO}, a world--action framework that separates \emph{what should change in the world} from \emph{how the robot should realize that change}. A predicted visual future is converted into an object-centric 3D representation containing a coarse rigid transformation $(R,t)$ and sparse contact or alignment references. A geometry-conditioned inverse dynamics model then maps this representation to short, closed-loop action chunks. Across five evaluated primitive settings, WECO achieves $82.5\%$ on Lift, $100\%$ on Move, $95.0\%$ on Rotate, and $100\%$ on a hinged TowelFold proxy across the tested conditions. Stack remains robust to appearance changes at the demonstrated layout but does not transfer to the tested geometric or spatial shifts. These preliminary results suggest that explicitly representing object-level world changes can provide a useful inductive bias for manipulation generalization, while also exposing remaining challenges in contact-rich composition.
Chat is not available.
Successful Page Load