Evaluating Physical Probes for Task Decisions under Hidden Dynamics
Abstract
Physical AI systems can collect force and proprioceptive measurements before acting, but an informative interaction does not necessarily improve the downstream decision. We present an evaluation protocol that separates task feasibility, available evidence, state estimation, and action planning under hidden mass, friction, damping, actuator gain, and delayed resistance. Across independently seeded 2,048-world MuJoCo studies, numerical planning from estimated state remains near 30%, while the same planner with true state completes every world. Neither eightfold more training data, a decision-aware objective, a different planner, 401 observations, nor direct maximum-likelihood fitting of the known simulator closes the gap. The direct fit adds 1.27 points (95% interval -0.44 to 2.98). A locked search over 35 feasible waveforms selects different probes for downstream completion and D-optimal parameter information. Task selection improves fresh-world completion by 3.71 points over the D-optimal multisine (1.90-5.57), below the registered five-point requirement, and remains 3.08 points behind the best fixed sweep. A finite-pool decision-conflict certificate distinguishes passive from active sensing but fails its registered association with decoder error; a post hoc reference-pool decoder points to sparse coverage of the continuous parameter space. On held-out TartanDrive 2.0 terrain, temporal history improves prediction by 23.53% over linear ARX but only 1.99% over instantaneous ExtraTrees. These results locate the main controlled-task loss before planning and show why parameter information, task-based selection, and strong fixed controls should be tested separately before a learned probing policy is credited.