Towards Semantically Diverse Skill Discovery through Contrastive Learning
Abstract
While deep reinforcement learning has achieved remarkable success, it remains constrained by the need for human-engineered reward functions for each different tasks, making it difficult to generalize across different scenarios. Unsupervised skill discovery methods seek to overcome this by allowing agents to intrinsically explore and learn diverse behaviors without external supervision. However, a significant challenge remains: because conventional methods often prioritize state-space diversity, the acquired skills are often semantically irrelevant to human observers. To address this limitation, we introduce a novel skill discovery method that leverages large language models to guide exploration toward learning semantically diverse skills. Our approach utilizes contrastive learning to align agent trajectories with corresponding language descriptions provided by a large language model, producing a semantic skill representation space. This skill representation is used to compute intrinsic rewards, guiding the agent to learn distinct, meaningful behaviors based on trajectory-level semantics. Empirical evaluations in MiniGrid environments demonstrate that our approach explicitly learns semantically relevant skills, such as retrieving the key or navigating to a locked door, that baseline algorithms often overlook. Furthermore, we also introduce a new evaluation methodology to measure semantic diversity, in which our algorithm achieves the highest score among the evaluated baselines.