Harness Optimization for Embodied Agents via Experience Traces
Abstract
Embodied agent performance is shaped by the capabilities of the underlying vision-language model and by the harness that structures interaction with the environment, including prompting, tool use, and failure recovery. Refining this harness can improve agent behavior without updating model weights, yet manual harness design remains labor-intensive and nontrivial. We introduce EMbodied Harness Optimization (EMHO), a framework that evolves the harness of a frozen embodied model from interaction experience under sparse environmental feedback. EMHO collects execution traces, identifies behavioral weaknesses, and evaluates harness revisions through sequential interaction. Across navigation and manipulation tasks, EMHO improves task performance and generalizes to held-out environments without updating the model. Qualitative analysis shows that evolved harnesses learn reusable recovery behaviors, including breaking navigation loops and verifying task states after failed actions.