When Is Visual Warning Physically Actionable? Controller-Conditioned Evaluation of Humanoid Hazard Avoidance
Nihar Mudigonda ⋅ Karn Kaura
Abstract
Physical scene information is useful for embodied decision-making only when it arrives in time and can be converted into improved physical outcomes by the downstream controller. We study this requirement in simulated humanoid hazard avoidance through a privileged-state, controller-conditioned evaluation. An onboard camera determines when warning becomes causally available; only then may privileged current geometry trigger and select one of two fixed responses executed by a learned whole-body controller. Across $480$ paired scenarios, we compare no intervention, post-contact reaction, hazard-independent conservative action, and selective pre-contact action. Among $120$ scenarios that collide without intervention, the privileged-state intervention prevents $42$ contacts, reduces mean peak force by $54.9\%$, and reduces maximum-link force-time severity by $70.8\%$. It produces one false trigger and no induced contact across $360$ controls. These gains carry costs: falls increase from $2$ to $14$, and goal successes decrease from $478$ to $466$. More importantly, aggregate performance conceals a large realization asymmetry: contact avoidance is $63.3\%$ for left-approach hazards and $6.7\%$ for right-approach hazards, while mirrored body-frame commands produce $0.15\,\mathrm{m}$ versus $1.17\,\mathrm{m}$ of stabilized lateral motion. This case study shows why physical-understanding methods should be evaluated through causal, controller-conditioned outcomes, not detection, predicted clearance, or aggregate safety measures alone.
Chat is not available.
Successful Page Load