Compute Allocation for Planners with Physically Misspecified World Models
Abstract
A model-based agent uses its world model as a simulator to plan inside, so the value of additional planning compute depends on how faithful the model's physics is. We study this dependence directly. Starting from a continuous pursuit-evasion task, we train ten world models that are identical in architecture, data volume, and training procedure, and differ only in a single physical constant of the simulator that generated their training data: drag, body radius, or integration timestep, each scaled over three severities. We then sweep the three compute parameters of a CEM model-predictive planner, imagination depth H, search width K, and refinement iterations N, against a control planner that uses the true dynamics on identical episodes. Three results emerge. First, each model has an interior optimal imagination depth, and the return on additional depth falls monotonically with the severity of the physical error, from +14.54 return units for quadrupling the horizon under correct physics to -12.77 under severe misspecification. Second, although the three parameters enter the rollout budget identically through KHN, they respond differently to misspecification: the return on width and on iterations also declines, but only depth becomes harmful, consistent with depth being the only parameter that increases how many times model error is fed back into itself. Third, at a fixed rollout budget the allocation itself matters: all ten models achieve higher return from three refinement iterations at depth 8 than from a single pass at depth 24, and at a smaller budget the preferred allocation depends on the model's fidelity. Physical fidelity therefore determines not only how well a planner can perform but how its compute should be allocated.