TimelineBench: Trajectory-Free Evaluation of Computer-Use Agents for Video Editing
Abstract
Computer-use agents are increasingly capable of operating professional creative software, yet benchmarks for video-editing agents remain limited. Existing evaluations typically assess rendered video quality or rely on intermediate milestones, leaving the resulting editable project state underexamined. We introduce \textbf{TimelineBench}, a benchmark of 100 video-editing tasks that evaluates computer-use agents through deterministic, trajectory-free comparison of structured project-state changes. Given a natural-language instruction and screenshots of an open-source video editor, an agent completes each task using only low-level mouse and keyboard actions. This design accepts alternative valid workflows and measures both binary task success and predicate-level partial progress. We validate every task through reference execution to ensure that the requested edits can be realized from the initial project and accepted by the evaluation conditions. Across five models and nine model--protocol configurations, the best configuration achieves 49.0\% task success, completes 64.8\% of requested conditions on average, and passes preservation checks in 92.0\% of tasks. Performance is substantially lower on operations that change clip timing or arrangement than on visual-property edits such as scaling, cropping, and opacity adjustment, and no open-weight configuration exceeds 10.0\% task success. These results show that reliably modifying editable video projects remains challenging even when agents make substantial partial progress.