PetTalk: A Latency-Aware Edge--Cloud System for Real-Time Affective Interaction
Abstract
Companion-animal vocalizations are brief and event-driven, making real-time interaction a joint problem of continuous sensing, selective cloud invocation, and response latency. We present PetTalk, a deployed edge-cloud system that closes the loop from companion-animal vocalization to spoken affective interaction. An always-on edge gate filters non-target audio before cloud-based affective-vocalization classification, response generation, and streaming speech synthesis. A pre-synthesized greeting provides early audible feedback while the complete LLM-TTS response continues. Across 40 stratified test clips, greeting playback onset is directly measured end to end at the speaker at 1.737 s at p50 and 3.003 s at p95. Complete-response onset is instead estimated from stage-wise percentiles at 2.86 s (p50) and 5.37 s (p95), with LLM generation dominating the latency budget. On a deterministic 90-minute household replay, the full edge gate reduces would-be cloud activations to 4.0% for cats and 3.1% for dogs, with 100% and 48% target recall, respectively. Supporting evaluation on an expert-annotated 1,346-clip set yields 95.7/94.0% classification accuracy for cats/dogs, while a convenience-sampled formative study (N=20) reports 9.0/10 latency acceptability. Together, these results support the feasibility of a practical latency-aware interaction loop while exposing the current dog-gating limitation.