Benchmarking Long-Lived Agents over Weeks of Replayed Reality
Abstract
We evaluate AI agent's capability to run unattended for days or weeks watching the streaming real-world environment and acting when needed. Unlike long-running agents that solve complicated coding or research tasks non-stop with a large number of steps, long-lived agents in our setup must decide when to wake, make meaningful progress, stay under a cost limit, and accumulate experiences over the long horizon. We evaluate agents over an evaluation suite of 6 tasks built from chronologically replayed real-world streams such as prediction markets and news archives, spanning weeks of world time. Tasks are divided along several axes, e.g., daily digest tasks that demand fixed polling or breakout news detection that requires vigilance and quick actions, all evaluated with deterministic metrics such as accuracies or delay of an alert. We evaluate both agent harnesses and base models and show how harness designs, such as wake trigger mechanisms and test-time improvement methods, jointly affect the mean and variance of agent performance.