When Objective Tuning Becomes Adaptation: A Leakage-Resistant Protocol for Robot Post-Training
Abstract
Post-training is a closed loop: evaluate a robot policy, diagnose failures, change an objective or update rule, and evaluate again. Unless this loop is partitioned carefully, information from test environments can enter policy selection and turn an apparent zero-shot result into manual adaptation. We illustrate the issue with a small, non-foundation-model case study in which a genetic algorithm trains a neural driving controller under readable distance, speed, and sensor objectives. On one 2D track with three runs per setting, a distance-dominant objective reduced logged evaluations by 35.7% and runtime by 70.4% relative to an all-zero pipeline control. Raising only the speed weight then increased evaluations by 16.7% and runtime by 208.0% relative to the best setting; exploratory sensor weighting also produced stalls and loops. The experiment cannot establish safety, transfer, or foundation model performance. We use it as a boundary case to propose a post-training protocol that logs objective provenance, separates development from held-out diagnosis, classifies failure and recovery, and labels every reported result by its adaptation budget. The same accounting is increasingly important when robot foundation models admit richer feedback and more powerful correction loops