Why Cross-Skeleton Retargeting Is Non-Identifiable: Structural Limits of Generative Motion Models
Abstract
Cross-skeleton motion generation trains generative models to carry action structure and motion intention from one body to another. Yet an action-consistent target motion admits two observationally compatible explanations: source-preserving transfer and target-action recovery. We show that this ambiguity is structural rather than incidental: under standard generative objectives, the source-conditioned retargeting map is non-identifiable in sparse heterogeneous motion domains. Unpaired marginal matching yields \emph{gauge non-identifiability}: a relative gauge between skeleton-specific latent spaces allows different source-conditioned maps to induce the same training evidence. Sparse paired supervision admits the complementary failure mode, \emph{conditional-mean degeneration}: ambiguous action-level pairings drive squared-error objectives toward a target-action prototype that is independent of the source clip's motion intention. To make the missing evidence observable, we introduce \emph{Source-Instance Fidelity} (SIF), a diagnostic that tests whether the source-side geometry remains visible in the generated target motions after the target skeleton and action are fixed. Under this diagnostic, standard action-level success across animal motion domains recurrently coincides with behavior at the \emph{source-blind floor}, while positive cases localize source-instance information to auxiliary motion-space constraints that narrow the observational equivalence class. The implication is that retargeting requires objectives and evaluations capable of identifying the source-conditioned map that a retargeting claim asserts. Project page.