TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development
Abstract
While LLMs excel at isolated coding tasks, their ability to perform autonomous, long-horizon machine learning (ML) development remains bottlenecked by poor strategic planning. Because current benchmarks rely on outcome-based metrics, they obscure process-level failures, making it impossible to diagnose where agent strategies diverge from human expertise. To solve this, we introduce TraceML, the first large-scale dataset of paired human and agent trajectories for long-horizon ML tasks. TraceML contains 4{,}672 Kaggle trajectories across 134 MLE-bench competitions, with 430 expert humans paired head-to-head with 207 LLM-agent runs on 7 of those competitions. Using a unified extraction schema to analyze step-by-step behavior, we reveal systematic gaps: agents execute local edits well but exhibit unstructured exploration, poor experiment prioritization, and weak long-term memory attribution compared to experts. Translating these diagnostic insights into action, we design a novel, planning-oriented agent harness based on human priors that significantly improves both process efficiency and final competitive performance.