Learning Symbolic Task Grounding for Temporally Structured Reinforcement Learning
Abstract
Many reinforcement learning tasks require objectives with temporal structure beyond a scalar reward. Reward Machines provide a principled representation of such objectives by tracking task progress based on symbolic propositions, but their use typically assumes a labelling function mapping observations to propositions. We study the setting where the reward machine is known but this symbolic task grounding is unavailable at deployment. We propose a world-model-based approach that learns this grounding from privileged proposition labels available during training. Building on DreamerV3, our method predicts propositions from latent world-model states and uses the given reward machine to generate rewards and track task progress during imagined rollouts, with its state conditioning policy learning. This provides an explicit interface between a learned world model and the symbolic task specification while requiring only raw observations at deployment. Experiments on temporally structured navigation tasks show substantial improvements over generic world-model and privileged-information baselines.