Four Planners, One Building: Shared-Task Benchmarking of Neural and Neurosymbolic World Models, and Three Ways It Misleads
Abstract
Planning models from the computational- neuroscience and neurosymbolic literatures are each evaluated on the task their authors designed, with the metric their authors chose. Cross-model claims are therefore mostly folklore. We put four such models on a single task: a generalized four-rooms building with a goal that relocates fifteen times. The models are a generative cognitive map learner (GCML), an active predictive coding agent (APC), a rollout-based meta-RL planner in the style of Jensen et al., and SSP-BO, a neurosymbolic Bayesian optimizer over Spatial Semantic Pointer embeddings. Two findings matter more than the ranking. First, path quality does not separate the models that work: GCML, APC and SSP-BO land at 1.18–1.65× the shortest path with standard deviations of 0.45–1.32, well inside each other’s spread, while their per-goal planning cost differs by four orders of magnitude (1.2 ms to 8.1 s). The informative axis is cost against quality, not quality alone. Second, three benchmark artifacts moved our numbers by more than the models did: a shared global RNG that changed one model’s score from 1.26 to 2.09 when an unrelated model was added to the evaluation loop; single-seed instability of comparable size; and a mis-scaled objective that made a working optimizer look 7.6× worse than optimal. We report these as results rather than footnotes, and we separately reproduce, inside our task, SSP-BO’s constant-time sample-selection claim (≈23 ms flat versus 220 ms and growing for a GP baseline at n ≈ 260).