Curation Metrics Are a Post-Training Decision: Auditing Trajectory Smoothness for Adapting Robot Foundation Models
Abstract
Trajectory smoothness has been proposed as a cheap, training-free signal for curating imitation-learning demonstrations, most recently by rinse, which introduces Spectral Arc Length (SAL) and a contact-aware Trajectory-Envelope Distance (TED) and reports downstream gains on robomimic. Recent audit work showed that generic curation metrics on a synthetic-defect LIBERO testbed largely exploit episode length rather than demonstration content. We test whether the specific published metrics survive the corresponding controls on real data. On robomimic can/mh, SAL separates better from worse operators with AUROC 0.952, but episode length alone achieves 0.957, and inside the tightest duration-matched band SAL falls to 0.628 while raw mean jerk, the baseline rinse argues against, rises to 0.896. On can/mg, where every episode is exactly 150 steps and length cannot carry information, SAL is at chance (0.454) and TED is inverted (0.376) against the task-success label, while jerk reaches 0.706. Training GMM behavior-cloning policies on each metric's curated subset (10 seeds per condition), curating by TED costs roughly a third of downstream success relative to random selection of the same size (0.240 vs. 0.360, p=0.0001, surviving Bonferroni correction); path length is directionally harmful with weaker support, while SAL and jerk are not statistically distinguishable from not curating at all. Inspecting individual trajectories shows why: on this task successful manipulation requires sustained high-deviation motion through the contact phase, while failure manifests as smooth low-amplitude hovering, so a smoothness metric measures commitment with the sign flipped. The length confound replicates on a second task: on square/mh, SAL correlates +0.96 with episode duration and its skill AUROC falls from 0.898 to 0.480 inside a duration-matched band, while jerk rises from 0.431 to 0.896. For post-training pipelines that filter, reweight, or select demonstrations before fine-tuning a pretrained policy, the practical consequence is direct: the curation step can silently discard the corrective, high-effort demonstrations that fine-tuning most needs, while keeping the smooth, low-information ones. Our results extend the detection-vs-curation decoupling to named, published metrics, on operator labels rather than injected defects, and to a benchmark where the length confound is excluded by construction rather than by truncation.