Tokenizer Choice Shapes Generalization in State-Centric Learning for Planning
Abstract
Planning has traditionally been addressed with symbolic search guided by domain models and heuristics. Learning offers a complementary way to exploit structure shared across related problems rather than solving each instance from scratch, making generalization to unseen instances a central challenge for learned planners. Recent work has explored planning mainly through action-centric models that predict actions or full plans from problem descriptions. A parallel line of work instead learns goal-conditioned transition models over states (state-centric learning) and recovers actions by matching predicted successor states to symbolic successors. In this setting, however, prior work has relied mainly on Weisfeiler-Leman (WL) representations, leaving the role of tokenizer choice unclear. We present a controlled study of tokenizer choice in state-centric learned planning, comparing WL, shortest-path, GraphBPE, SimHash, and a deterministic baseline within a common pipeline. Across varied benchmark domains and multiple planner configurations, WL performs well overall, but no tokenizer is best everywhere: the leading representation changes by domain, and tokenizer differences are larger on extrapolation rather than on interpolation tasks. Across the studied configurations, tokenizer choice is the dominant factor in held-out generalization: it explains 57% of performance variance, more than predictor architecture or any other pipeline factor, and its effect is substantially amplified on out-of-distribution problems relative to in-distribution ones. We conclude that representation choice is a strong determinant of generalization in state-centric learned planning and that tokenizer choice is a first-order modeling decision rather than a fixed preprocessing step. Code: https://tinyurl.com/4hre257u