Archetypal Sparse Transcoders Find Motion Features and Physical-Hazard Detectors in V-JEPA 2
Abstract
As embodied systems are deployed in increasingly complex environments, it becomes necessary to monitor their physical safety from the inside. Sparse autoencoder (SAE) feature patterns are known to signal out-of-distribution operation and impending hallucination in vision and language transformers; we ask whether analogous probes can be built for the self-supervised video encoders now entering the physical-AI stack. We introduce an archetypal sparse transcoder for V-JEPA 2: a sparse decomposition that combines the layer-to-layer prediction objective of transcoders with the convex-hull decoder constraint of archetypal SAEs. Trained on the block-10 FFN of a frozen V-JEPA 2.1 ViT-B/16 using Something-Something v2, the transcoder reaches held-out explained variance 0.479 at L0 = 59, against a random-initialization control at −1.4×104. A label-distribution protocol over the 174 SSv2 motion templates surfaces features that track a common motion primitive across different objects, and a physical-hazard probe finds a statistically enriched sub-population that fires selectively on falls, drops, and collisions, 5.7× rarer in the untrained control. Over eight million held-out tokens per feature, these labels hold far beyond the top-activating clips (7.9× to 64× per-token enrichment), and preliminary multi-source training raises the dictionary’s alive fraction from 4.5% to 60%.