EvoEmbody: Self-Evolving Embodied AI Agents in Interactive Text Environments
Abstract
Self-improving embodied agents often rely on large proprietary models, risk con- text exhaustion, or evaluate updates on overlapping task pools. We introduce EvoEmbody, a self-evolving framework using frozen open-weight language models (≤31B) without gradient updates. EvoEmbody isolates error mining on a dedi- cated training set from candidate validation and uses on-demand recursive retrieval to process diagnostic traces without context overflow. On ALFWorld, EvoEm- body substantially improves held-out, out-of-distribution success rates: Qwen3.8- 27B/Qwen3.5-4B advances from 56.0% to 93.3%, and Gemma-4-31B/Gemma-4- 26B-A4b increases from 64.9% to 79.1%. Disjoint training-pool mining signif- icantly outperforms validation-pool mining (93.3% vs. 77.6%), as meta-agents autonomously resolve core bottlenecks in inventory tracking, state-conditioned matching, and spatial coordination.