DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning
Suyog Khanal ⋅ Arun Kumar A V ⋅ Santu Rana
Abstract
Learning embodied skills from limited expert supervision requires more than deciding when an expert should intervene; it also requires deciding what the available supervision should address. Interactive imitation learning provides a natural setting for this problem, as autonomous practice exposes deficiencies in the current policy while expert demonstrations remain costly. We introduce **DISEIL** (**D**emonstration d**I**stillation for **S**ample-**E**fficient **I**mitation **L**earning), which combines structured representations of policy failures with language-model reasoning to determine where subsequent demonstrations should be targeted. DISEIL identifies recurring deficiencies across failed episodes, groups related failures using geometric descriptors, and reasons over observations and trajectories from a selected failure mode to synthesize a demonstration starting configuration intended to address its shared deficiency. Explicit task and embodiment constraints ensure that the prescribed configuration is feasible, while the learned policy remains solely responsible for producing robot actions. Across $5$ manipulation tasks with both state and image observations, yielding $10$ experimental settings, DISEIL achieves the highest mean held-out success rate under the same demonstration budget in every setting. These results show that structuring and reasoning over failure experience can make limited supervision more effective, and provide an initial step toward embodied learners that reason not only about how to act, but also about what they still need to learn.
Chat is not available.
Successful Page Load