Planning Guided Option Critic
Abstract
Symbolic skill specifications give reinforcement learning agents structure for long-horizon tasks, but executing them in a fixed order prevents the agent from adapting execution to the state it encounters. We introduce Planning Guided Option Critic (PGOC), which couples symbolically specified skills with a learned policy over options. Skill preconditions and a continued-need predicate restrict which options are admissible at each decision point, and bounded plan-progress rewards direct the controller toward outstanding requirements. Each skill is trained on its local reward plus the controller's continuation value, so a completed skill is credited for the progress it enables rather than for the next skill's objective. On the Gnomish Mines progression in Craftax, PGOC with recurrent backbones reaches the final milestone in 43.1-43.7\% of held-out episodes, versus 7.7-9.5\% for SCALAR's prescribed skill sequence and at most 2.6\% for PPO. Within PGOC, ablations test the option masks and boundary target: removing prerequisite masking eliminates late-task progress, removing continued-need masking diverts practice toward already-mastered skills, and grading skills by the next skill's value in place of the controller's collapses later milestones at matched training frames.