Autonomous Research for Open-Ended Problems: A Case Study on Telecom Ticket Retrieval
Abstract
Recent breakthroughs in LLM-based systems and their abilities in problem solving and coding has allowed progress in the AI for Science paradigm, replacing human roles in machine learning (ML) research in part or even in full. However, while several frameworks of fully autonomous end-to-end ML research have been proposed, successful implementation of them are often limited to problems with narrow search spaces, like language modeling or biomedical ML benchmarks. In this paper, we explore how autonomous research frameworks can be adapted to solve open-ended, insdustry-grade ML problems, by considering a case study: telecom ticket retrieval, an open ended task with degrees of freedom in representation, architecture, and training data generation. We discover that autonomous research for open-ended problems with commercial and open-source agents shows both promise and limitations. Our best discovered single-model system (R@1 0.34) outperform a human-designed ensemble (R@1 0.25) yet suffer from operational overhead and fail to explore creative solutions like data augmentation, ensembling, or re-ranking, which the human-developed state-of-the-art system incorporates (R@1 0.38).