Exact Verifiers Refine Policies, Not World Models: Auditing Agent Improvement in Physical Control
Abstract
Can exact action scores guarantee that verifier-guided training produces a better agent? We examine this question in building control, where a simulator provides environment-grounded verification. We first audit a Twin Delayed Deep Deterministic Policy Gradient (TD3) critic on the same-state action comparisons used by a group-relative update. Although the critic is highly accurate along its reference trajectory, it often misranks actions sampled by the language policy. We therefore replace it with deterministic simulator rollouts under a fixed continuation policy. Even with exact scores, the updated agent does not improve reliably in sampled actions or closed-loop control. We show formally that an exact scalar return does not identify a unique action-conditioned transition model, and find that outcome-only training does not improve the model’s predictions of action effects. A separate experiment provides the complementary case: when the base model can already predict action effects and generate useful actions, training improves its reasoning and control performance. Together, these results distinguish policy refinement from world-model learning. Exact verifiers can refine behavior already available to the policy, but they do not teach the model how its actions change the environment.