Reset Dependence Is Embodied: Training Protocol Changes Which Morphologies Appear Optimal
Abstract
In a 100-body robot co-design benchmark, the height threshold that resets a falling robot changes which bodies appear optimal. We train PPO controllers for DERL-evolved UNIMAL morphologies under Standard (τ = 0.20), Strict (τ = 0.50), and Reset-free (τ = 0) protocols, then evaluate every policy under a common protocol. Strict and Reset-free training share only 5 of their top-10 bodies and have low rank agreement (Spearman ρ = 0.21). The effect is family-specific, not a uniform penalty: averaging per-body retention ratios, Reset-free training reduces Floor bodies' per-step reward by 45% relative to Standard training but reduces variable-terrain bodies by 16%. Under Strict, the losses are 22% for Floor and 34% for variable-terrain bodies, while multi-task bodies are most stable across protocols. Blocked ANOVA and morphology-level permutation tests support the protocol-by-family interaction. A SAC validation panel planned before analysis finds the same Strict-versus-Reset-free family ordering under a different optimizer. Selecting bodies by their minimum reward across the three training protocols recovers ≥91% of every single-protocol oracle's top-k average reward (k ∈ {5, 10, 20}; all bootstrap lower bounds ≥84.3%). Co-design benchmarks should state the reset rule, audit ranking sensitivity, and report protocol-robust selections when choosing bodies.