HiMA-Ecom: A Long-Horizon Memory Benchmark for Hierarchical Multi-Agent E-commerce Assistants
Abstract
Long-horizon agentic systems become long-context systems by accumulation rather than through a single long document, since every step reassembles a prompt from role instructions, tool and sub-agent descriptions, recalled memories, dialogue history, and previous tool outputs, separately for every agent. This paper studies how such accumulated context can be trained, evaluated, and managed in hierarchical multi-agent assistants. We present HiMA-Ecom, a hierarchical multi-agent benchmark for e-commerce assistants with 22.8K instances, of which 17.7K carry memory entries in context, pairing agent-specific supervised fine-tuning data with system-level supervision for end-to-end multi-agent reinforcement learning. Building upon it, HiMA-R1 keeps the context budget of a trajectory under control through Variance-Reduction Group Relative Policy Optimization (VR-GRPO), which samples along an initial trajectory so that sampling cost grows additively with trajectory length and updates the agents with the largest reward variance, together with an efficiency reward on trajectory length and a memory evolution mechanism that repurposes GRPO rewards as cost-free signals for what stays in context. Agents built upon 3B and 7B open-source models match DeepSeek-R1 and surpass DeepSeek-V3 by 6\% on average, while reward-driven memory reduces the update steps needed for convergence by roughly 55\%.