PRAXIS over Theoria: Self-Exposure in Place of Scale for Efficient Manipulation
Abstract
Closed-loop manipulation policies usually buy robustness to distribution shift with large pretrained backbones and iterative generative decoders. By asking how much of that robustness can come instead from the training signal, this paper introduces Self-Exposure with Aligned Labels (SEAL), which rolls the current policy out in simulation and supervises every state with the demonstration's actions, with or without an added analytic correction. SEAL trains PRAXIS, a task-conditioned vision-action policy learned from scratch with no language model, no pretrained backbone and no language input, at 7.56M parameters and 0.831 mean success on LIBERO, against 0.821 for CT-VAM, a language-free policy at nine times the size. SEAL needs no interactive expert to relabel the visited states. That requirement, is what has limited DAgger-style correction, and on a fixed task repertoire the demonstration can supply the label.