Skill-Level Effects in Behavioral Cloning: When Low-Skill Data Improves Performance
Abstract
Behavioral cloning (BC), which trains models from offline demonstrations, is a common approach in reinforcement learning settings. Prior work argues that BC requires expert demonstrations and performs poorly when trained on low-skill data. We challenge this assumption by showing that, in certain regimes, training on low-skill data can yield models that outperform those trained on high-skill data. Because low-skill data is often cheaper and more easily acquirable, this finding has important practical implications. To explain this result, we introduce the notion of fragility, characterizing how a policy's reward degrades under errors, and provide theoretical insights on how fragility can predict low-skill outperformance. We test our approach in a synthetic environment and MuJoCo and validate it using human data from chess and racing. Motivated by these findings, we connect our results to curriculum learning by structuring training according to demonstrator skill, rather than task difficulty as in standard curriculum design, and show that such skill-based curricula can improve performance relative to standard BC approaches.