Environment-Grounded Verification of Humanoid Hazard-Avoidance Improvements
Abstract
Agent modifications are often judged using scalar success metrics or intermediate signals that may not reflect realized behavior. We study this verification problem in simulated humanoid hazard avoidance. We treat the complete warning-to-execution stack as an embodied agent and ask whether adding a warning-timed, geometry-informed intervention actually improves it. An onboard camera determines the first time a hazard could realistically be seen; from that point on, a privileged rule with perfect current geometry may choose between two fixed evasive responses, both executed by the same learned whole-body controller. Across 480 paired scenarios, the full informed intervention prevents 42 of 120 baseline contacts, reduces mean peak force by 54.9%, and reduces maximum-link force-time severity by 70.8%. Yet these favorable signals give an incomplete verification result: contact avoidance is 63.3% for left-approach hazards and 6.7% for right-approach hazards, selected predicted clearance does not reliably track realized collision, and 69 of 78 residual contacts occur during recovery or stabilization. The intervention also increases falls from 2 to 14 and decreases goal successes from 478 to 466. The results apply to the complete intervention rather than the marginal effect of geometry alone. As a case study in environment-grounded verification, the study shows why agent improvements should be evaluated through execution using heterogeneous outcomes, subgroup results, and a horizon that includes recovery.