World Models for Temporally Structured Tasks with Unknown Symbol Grounding
Abstract
Many reinforcement learning tasks require objectives with temporal structure beyond a scalar reward. Reward machines provide a principled representation of such objectives by tracking task progress based on symbolic propositions, but their use typically assumes a labelling function mapping observations to propositions. We study the setting where the reward machine is known but this grounding is unavailable at deployment. We propose a world-model-based approach that learns the grounding from privileged proposition labels during training. Building on DreamerV3, our method predicts propositions directly from latent world-model states and uses the given reward machine to generate rewards and track task progress during imagined rollouts, with its state conditioning policy learning. This integrates the known task structure into both world-model learning and imagination while requiring only raw observations at deployment. Experiments on temporally structured navigation tasks show substantial improvements over generic world-model and privileged-information baselines.