FOCUS: First-Order Control Using Optimal Transport Supervision for Imitation from Observations
Abstract
Imitation learning methods driven by trajectory-matching objectives typically rely on temporally aligned expert and agent trajectories and require full-horizon backpropagation when paired with differentiable simulation. In contact-rich control, this often produces high-variance, brittle gradients, and can fail under even modest temporal misalignment between expert demonstrations and agent rollouts. We propose FOCUS, a first-order imitation learning from observations method that combines differentiable physics with optimal transport (OT)-based local alignment. FOCUS trains from short-horizon agent rollouts by matching them to multiple state-only expert segments, yielding stable local imitation objectives without requiring temporally aligned trajectories or full-horizon backpropagation. These local OT objectives are coupled with critic bootstrapping, allowing short-horizon differentiable supervision to propagate across longer-horizon behaviors. Evaluated on contact-rich continuous-control benchmarks in DFlex, including a challenging musculoskeletal humanoid, FOCUS improves sample efficiency and training stability over adversarial imitation-from-observation methods and prior differentiable imitation baselines.