Open-Ended Autonomous Model Research for Recommendation Systems
Ziqi Chen ⋅ Mohamed Hammad
Abstract
We present an autonomous research system for open-ended model development in recommendation systems. At its core is a batch-wise, evidence-conditioned research loop: a central Researcher constructs hypothesis portfolios spanning new and promising prior directions, while independent Coder subagents implement and evaluate selected hypotheses in parallel. The system separates stochastic scientific reasoning from deterministic experiment execution and externalizes long-horizon state into persistent stores for hypotheses and analyses, experiment trials, and debugging knowledge. After the human-defined task contract is fixed, no human proposes the candidate hypotheses reported as qualifying results; without the agent-generated interventions, the corresponding results reduce to the human-provided reference configurations. Across two public recommendation benchmarks, the system develops non-monotonic research trajectories spanning architecture, objectives, representations, and optimization. On KuaiRand-1K, the strongest reviewed configuration improves the task-defined aggregate score from $0.1570$ to $0.2041$ (+30.0\%) and MRR from $0.0675$ to $0.0850$. Under our full-data KuaiRec protocol, the strongest reviewed configuration reduces MAE from $4.6621$ to $4.4517$ (-4.51\%) and increases XAUC from $0.5532$ to $0.6094$ (+0.0562) relative to the EGMN reference trained under the same protocol. We additionally report qualitative transfer evidence from a production-scale workload. Beyond final performance, we analyze the research process itself. A post-hoc taxonomy of 500 stored hypotheses characterizes task-specific differences in research emphasis and multi-family composition, while retrospective human audits uncover operationally successful but scientifically invalid trials. These failures expose a gap between reliable experiment execution and valid scientific evidence.
Chat is not available.
Successful Page Load