The Puppet Test: Toward a Nonverbal Interaction Benchmark for Human-Robot Interaction
Abstract
Large language models are now able to pass the original, text-based Turing Test. As intelligence enters the embodied realm, a new test is needed to benchmark interaction beyond speech. We propose the Puppet Test, a nonverbal Turing Test for embodied agents, designed to assess whether autonomous robot behavior can be distinguished from human-controlled behavior during live, speech-free interaction. The test asks human judges to interact nonverbally with the same embodied, anthropomorphic agent in two consecutive sessions, one autonomous and one controlled by a hidden human puppeteer, and then identify the condition in each session. Unlike prior nonverbal Turing Tests that evaluate individual behaviors such as gaze or gesture, or rely on prerecorded interactions, the Puppet Test provides a reproducible study design and pass criterion based on whether judges distinguish autonomous behavior as human control beyond chance. We report a first baseline implementation of an autonomous nonverbal interaction system that combines continuous mirroring, responses to gestures and facial expressions, and a vision-language model controller. The system was judged as a human in 8 of the 30 cases, and therefore the challenge remains open for future systems to meet the pass criterion. Qualitative analysis further revealed the nonverbal cues and reasoning participants used to make their judgments. Finally, we provide recommendations for future studies using the Puppet Test. Together, the test, teleoperation interface and reference autonomous implementation on the widely available Reachy Mini provide a practical path towards benchmarking nonverbal interaction in embodied AI and support comparison across future systems.