TGPO: Temporal Grounded Policy Optimization for Signal Temporal Logic Tasks
Abstract
Learning control policies for complex, long-horizon tasks is a central challenge in autonomous systems. Signal Temporal Logic (STL) offers an expressive language for specifying such tasks, but its non-Markovian nature and inherent sparse reward make it hard for standard Reinforcement Learning (RL) to solve. Prior RL approaches focus only on limited STL fragments or use robustness scores as sparse rewards. Our new method TGPO (Temporal Grounded Policy Optimization) decomposes STL into timed subgoals and invariant constraints and tackles the problem in a hierarchical fashion. Its high-level module proposes time allocations for subgoals, and the low-level time-conditioned policy learns to achieve the sequenced subgoals using a dense, stage-wise reward. In inference, we use the critic to efficiently sample time allocations and select the most promising assignment for the policy to rollout. We evaluate in five environments, from low-dimensional navigation to manipulation, drone, and quadrupedal locomotion, and TGPO significantly outperforms baselines (especially for high-dimensional and long-horizon cases), with 31.6% higher success rate compared to the best method.