LocalLip Live: The Face of a Conversational Avatar, on a Fanless Laptop, Offline
Abstract
We run an interactive avatar service in production: a visitor converses with an on-screen person chosen in advance, on kiosks, on the web and on mobile. Of that chain the face is the stage the operator has to run itself, and the one that has kept a GPU behind every stream. We demonstrate that stage, and only that stage, on a fanless M4 MacBook Air with the network off: the laptop generates the video at 33 frames per second, 28.3 ms per frame, every one of 230 frames inside the 30-fps period---faster than it plays, while an unoptimised port of a 1.27B diffusion model on the same laptop needs 56 s for one frame. Personalizing the face removes that factor: the avatar is fine-tuned once, beforehand, from a few minutes of video. Visitors put the laptop in airplane mode themselves, and see the same model fail on the same identity once the personalization is removed.