MLE Agents: Autonomous ML Engineering in Production Recommendation Systems
Abstract
The typical workflow for improving large-scale ranking and recommendation models is slow and tedious: an engineer forms a hypothesis, implements it as a code change, trains a model, and evaluates it both offline and online, often relying on their own intuition and experience along the way. Autonomous LLM agents promise to speed this up by running the loop on their own, but it is still unclear whether they can do this in real industry-scale production systems rather than on benchmarks or small proxy tasks. We build human-on-the-loop machine learning engineer (MLE) agents that propose, implement, train, and evaluate model changes on production search rankers serving over a billion requests a day. Engineers retain control over objective design, search direction, final-model review, and online deployment. Each agent writes a model change as code, launches a training job, reads the offline evaluation, and decides whether to keep it. We identify three obstacles that make this industry setting challenging: feedback takes hours per experiment, offline metrics do not always predict online impact, and improvements saturate after the early wins. We address these with parallel training runs, an objective calibrated to match online impact, and careful context engineering (a "what didn't work" blocklist, strict output formats, and an anti-staleness rule). In one campaign, the system ran 1,097 completed candidate experiments and accepted 325, raising the win rate from 15.6% to 29.6% of candidate runs. Its best model launched to production, increasing click-through rate by 0.70% and search volume by 0.23% in an online A/B test on approximately 10 million users. Several of these improvements were novel architectural changes the team had never tried -- mechanisms drawn from the broader industry and literature that had not previously been considered for this stack -- suggesting that MLE agents can explore regions of the design space that human intuition overlooks. Notably, we find that the design of the harness around the model matters significantly: intuition suggests that a stronger underlying model is enough, but in our experience, the design of the loop around it is just as critical. More broadly, our work shows that autonomous agents can do real machine-learning engineering in high-stakes production systems.