Robo-Cortex: Test-Time Continual Strategy Learning for Embodied Agents
Abstract
Test-time continual learning (TTCL) is often associated with updating model parameters, yet an embodied agent can also learn persistently by revising the strategy that governs its future decisions. We study this strategy-layer form of TTCL. Existing memory and reflection systems retain prior trajectories, but often leave reusable decision principles fixed, a failure mode we call strategic stasis. We introduce Robo-Cortex, which keeps its foundation models frozen while converting deployment experience into a persistent, versioned bank of navigation heuristics. After each interaction round, the agent diagnoses successes and failures, induces candidate rules, and activates only rules supported by sufficient confidence and recurring evidence. The resulting strategy is then redeployed through closed-loop planning. Across image-goal navigation, active recognition, and active embodied question answering, Robo-Cortex improves over non-evolving and reflection-based agents. On image-goal navigation, success rises from 36.62% to 51.39% across five evolution rounds. Holding retrieval fixed while updating strategy outperforms updating memory alone, and transferred self-induced rules reach 48.61% success versus 38.89% for human-written guidance. These results support persistent strategy revision as a practical TTCL mechanism for embodied agents.