Asymmetric On-Policy Distillation Learns In-Context Exploration
Sriyash Poddar ⋅ Yanda Bao ⋅ Jacob Krantz ⋅ Matthew Chang ⋅ Xavier Puig ⋅ Roozbeh Mottaghi ⋅ Natasha Jaques ⋅ Abhishek Gupta
Abstract
Sequential decision-making under partial observability is difficult because agents must explore to infer hidden, task-relevant information before acting effectively. Learning such information-gathering behaviors \emph{tabula rasa} via online reinforcement learning is sample-inefficient, whereas offline imitation fails to learn exploratory behaviors due to insufficient coverage of deployment-time histories. To address this, we develop an online imitation-based method that learns to explore and adapt during deployment. We focus specifically on the setting of asymmetric student-teacher learning under partial observability using oracle information from clairvoyant experts. Our method, ASTEROID, is based on the key idea that history-conditioned student policies can learn to explore through online asymmetric distillation: the student first collects on-policy context under partial observability, and the privileged expert then provides action labels conditioned on that context. Iteratively, as the context grows, the student learns to use past observations to infer hidden task information and act according to the inferred latent state, yielding exploration similar to efficient Bayesian posterior sampling. Across a diverse range of eleven different simulated environments, including discrete exploration benchmarks, robotic manipulation and navigation environments, our method achieves $\sim$2$\times$ higher returns under fixed exploration budgets and uses up to $\sim$100$\times$ fewer interactions than baselines. Finally, ASTEROID transfers policies trained entirely in simulation to vision-denied real-world manipulation, exhibiting efficient exploration and improving success by $40\%$ over baselines. Overall, our approach provides a simple yet surprisingly effective framework to learn efficient test-time exploration from privileged supervision.
Chat is not available.
Successful Page Load