Probe--Exploit: When and How to Recover Hidden Symmetry Breaker in Equivariant RL
Yibin Fan ⋅ Qi Zhang
Abstract
An unobserved symmetry breaker can remain fixed within each trial while its transformed versions are resampled across trials, leaving the marginalized task exactly symmetric. We study when an equivariant RL agent should recover such a breaker and how the recovered information should be used. We first compute a pre-training upper bound $U$ on the potential benefit of observing the breaker for free: $U=0$ rules out such a benefit within our policy comparison, whereas $U>0$ only leaves it possible. Probe--Exploit then identifies the active breaker and conditions a jointly equivariant policy on it. Across six breaker designs, we evaluate two contrasting cases end to end. Under a fixed 150,000-interaction budget, Probe--Exploit yields success-rate gaps of $+0.270$ on a larger-$U$ obstacle task and $-0.230$ on a small-positive-$U$ force field, relative to equivariant RL without recovery. These results support $U$ as a pre-training screen and show that recovering the breaker alone does not guarantee a learning benefit.
Chat is not available.
Successful Page Load