Agentic Data Engineering for LLMs: System Design and a Controlled Empirical Study
Abstract
AI-for-AI (AI4AI), which uses AI systems to improve the data, training algorithms, and feedback used to develop other AI models, is a promising and rapidly growing direction. Yet two questions remain open: can an AI4AI system propose strategies that improve a target model on evaluation data not used in the research loop, and can it iteratively refine these strategies using cumulative feedback? To study these questions under a controlled research process, we introduce Agentic Data Engineering (ADE), a system for LLM data engineering. In ADE, agents drive the research process: they propose strategies, choose what to explore, and refine their strategies using experiment feedback. A fixed protocol governs experiment execution and in-loop evaluation, while keeping held-out evaluation outside the research loop. We study ADE on data selection and reward design. Across in-loop-selected strategies, the largest absolute held-out gain over the generic baseline is 13.34% on AIME25 Pass@32. Across four held-out benchmarks, absolute macro-averaged Pass@K gains reach 5.85%, 2.72%, and 2.97% for Math SFT, Code SFT, and Math RFT, respectively. Reward design improves coverage under repeated sampling despite limited gains in average correctness. Research trajectories show successive in-loop improvements and evidence-guided strategy revision and combination, with returns on additional research capacity varying across tasks. Together, our study shows how a controlled research process can separately assess whether AI4AI discovers strategies that improve performance on unseen data and whether cumulative feedback leads to progressively better strategies.