Same Policy, Different Artifacts: Paired Task Audits for On-Device Neural Control
Abstract
A model labeled “INT8” can run as quite different controllers on a device. Posttraining quantization sets numerical ranges from a few representative states, and build records can omit which states were used and how the compiler used them. Holding one policy, the Hailo-8 toolchain path, and 1,024 paired task starts fixed while changing only which states calibrate the model yields 96.6% versus 21.4% terminal success and a return spread equivalent to a difference of about 30 cm in mean goal distance. Across 70 deployment configurations, every device-arm control step runs on a physical Edge TPU or Hailo-8 accelerator in simulated tasks. The effect recurs prospectively in 7/8 independently trained FetchReach policies, and outcome-informed checks at compiler optimization levels 1–3 retain it. On the tested path, the compiler reports 64 calibration rows without their identities, and at levels 0 and 1 builds given 256 rows match builds given only the first 64 in every audited action and task outcome, so reordering the same rows changes which rows set the ranges. Neither fixed-state action error nor our local diagnostic substitutes for complete-task evidence. We use deployment instance to denote the exact executable bound to its runtime, device, runner, and I/O processing. These findings motivate deployment-instance-specific task audits and provenance records; application-contract decisions are illustrated retrospectively. Judge a compiled artifact by paired task outcomes for its exact deployment instance and application contract, not by a checkpoint, precision label, or local proxy.