Real-Time Multimodal Conversational AI (ReMuCAI)
Akash Gupta ⋅ Vyas Raina
Abstract
Pose-guided human image animation recreates the motion of a reference video with a new subject appearance, conditioned on a driving \textit{pose sequence} and a small budget of appearance reference frames, which we call \emph{anchors}. Each anchor is assigned to a specific timestamp, where the video generator must match that frame; it then synthesizes the full clip subject to both these timestamped anchors and the pose sequence. Each anchor costs one image-model call, so their number is budgeted. The video generator is far more expensive, at a real-time factor of roughly eleven (30\,fps at $1024\times1920$ on one B200 GPU). Selecting anchor positions is important: placing anchors in the low-error regions of a clip's pose-error profile (error between reference and generated video) increases the mean error by $8.6\%$ relative to uniform placement, and the penalty grows with motion difficulty. Identifying those regions by measurement costs two video-generation calls. We instead predict the per-frame error profile before video generation with a 254k-parameter head on the generator's 319M frozen VAE encoder. On unseen, out-of-distribution videos it attains a Spearman correlation of $0.638$. The frames it selects as lowest-error sit at the 18th percentile of realised error, against the 14th when selecting on the measured profile. The ranking transfers without retraining to AIST++ and HyperMotionX. An agentic planner can therefore exclude wasted anchor positions before any generation, and distribute a fixed budget across clips, capturing $68\% \pm 28\%$ (mean $\pm$ s.d.\ across budget levels) of an oracle allocator's gain.
Chat is not available.
Successful Page Load