TimeWarp: Evaluating Web Agents by Revisiting the Past
Farhan Ishmam ⋅ Kenneth Marino
Abstract
As web agents close the gap with humans on benchmarks, it raises the question: *Do today's agents perform just as well on tomorrow's web?* We introduce TimeWarp, a benchmark that emulates the evolving web. TimeWarp consists of three web environments, each with six UI versions spanning UI design and frontend code from different eras of the internet. We pair TimeWarp with a set of complex, realistic tasks covering different forms of web navigation. Our experiments reveal vulnerabilities of web agents, especially visual ones, to changes and the limitations of behavior cloning (BC) on complex trajectories from a single version. To address this, we propose TimeTraj, a simple yet effective algorithm that uses plan distillation to collect trajectories across multiple versions. By training agents on teacher rollouts using our BC-variant, we achieve substantial performance gains: $20.4$ %$\rightarrow37.7$% for Qwen-3 4B and $0$ % $\rightarrow27.0$% for Llama-3.1 8B models. Our work helps study generalization across web designs and enables a new paradigm for collecting plans rather than trajectories to improve the robustness of web agents.
Chat is not available.
Successful Page Load