Evaluating Action Representations for 3D Orientation in Model-Based Reinforcement Learning
Abstract
Many robotic manipulation tasks (insertion, door opening, tool use) require controlling orientation, not just position. However, choosing how to parameterize SO(3) is itself a design problem: every representation has a drawback, from discontinuity to over-parameterization. These trade-offs have been studied in supervised learning and model-free reinforcement learning (RL). But they remain largely unexamined in model-based RL (MBRL), where a learned world model must represent actions as conditioning input to its dynamics function. We present the first evaluation of SO(3) action representations under a model-based learner, comparing quaternions and tangent vectors in global and delta mode with TD-MPC2. We find the model-free results do not transfer cleanly to model-based results across our experiments. We recommend that rotation parameterization be re-validated per learner class rather than inherited from model-free results.