Verifiable Step-Wise Planning: Cost-to-Go Reward Design for Language Model Planners
Abstract
Language models can do sequential planning, but reinforcement learning with outcome-only rewards provides limited information about which intermediate actions contribute to a successful plan. In verifiable planning environments, the exact optimal cost-to-go for each state can be found through symbolic computation. In this work, we propose an exact cost-to-go reward for GRPO-trained small language model planners based on differences in optimal remaining cost, which allows for verifiable step-wise credit. We compare our proposed reward design to an outcome-only reward design, on Qwen2.5-1.5B-Instruct in weighted gridworld planning tasks under matched training conditions. In greedy evaluation, A4 achieved a 1.2\% success rate on 2,000 held-out instances, compared with 0.4\% for the base model. These findings imply that exact and verifiable step-wise feedback from the environment can provide a useful training signal compared with outcome-only supervision without sacrificing verifiability. Our code will be made publicly available upon acceptance of the paper.