Invariance without Equivariance: Style-Action Coupling in Text-to-Motion
Abstract
Add "tiredly'' to a motion caption. The manner should change and the action should survive. Everything built on these models to edit, retrieve or steer assumes that separation already holds in the representation, and the scores the field reports never test it. We ask what transformation styling actually applies to a model's arrangement of actions, and separate two properties it could have. Relational invariance holds when styling preserves the similarity structure among actions, an orthogonal transformation of the configuration; directional equivariance holds when the style displacement points the same way at every action, a translation. Neither implies the other, and only the second is what a steering vector can use. We measure both across four generators, three text-motion alignment encoders and the CLIP text input, on 24 actions crossed with 13 caption-mined manner modifiers, read against permutation nulls and tested interventionally. The two come apart. Relational structure varies widely across architectures, while the style direction collapses in all four generators, further than heading, a non-style attribute measured the same way. Latent diffusion is the sharpest case, matching the input's relational structure while keeping almost none of its direction, though a causal test recovers a weak residual there. All four condition on a frozen CLIP text encoder, so they are handed that direction and lose it themselves; the alignment encoders lose most of it too. Where it can be tested, style adds correctly at each action but the same vector does not carry to another. Relational invariance is therefore not sufficient for steering: a similarity-only diagnostic will certify a representation that latent arithmetic cannot use. The diagnostic pipeline and prompt grids will be made publicly available upon acceptance.