Do LLMs Have Research Taste? An Analysis Using BóLèBench
Abstract
Large language models (LLMs) are increasingly used for recursive self-improvement, generating and implementing new research ideas that advance their own capabilities. Since compute is finite, a central question emerges: which ideas should be pursued before any results exist? For human researchers this is a matter of taste, honed by years of ideas succeeding and failing. LLM judges now make this call and allocate real compute, yet their taste has never been measured directly, since existing evaluations grade against publication records, which hide failures and reward memorization. We introduce BoleBench, a benchmark built on a simple inversion: use hindsight to grade foresight. We start from parallel rollouts, where different agents attempt to improve an expert baseline on a fixed set of ML research tasks under matched compute budgets, yielding many scored attempts per task. A judge reads an abridged version of each trace, with results removed, and predicts which will score higher. The executed outcomes serve as ground truth, so failures stay visible without human annotation. BoleBench is a pipeline for turning executed attempts into graded test questions, supporting five kinds of research decisions graded at four levels of informativeness, grading each judge across multiple criteria. New question banks are built from post-cutoff corpora as models retrain. Our findings indicate that current LLM judges fare poorly. A simple "pick-the-bolder-attempt" heuristic beats every frontier judge we test, and no judge beats chance on dark-horse pairs, where the bolder attempt loses. Judges across four model families miss the same pairs at roughly 135 times the independent-error rate, and a bad judge selecting among candidates is worse than picking at random. Finally, we reveal a "model name bias": revealing an attempt's authoring model shifts rankings toward prestigious frontier models, regardless of quality. We release BoleBench: 3,863 blinded pairs from four corpora, among them 1,842 pairs from 271 executed agent attempts on ten ML-research tasks and 1,701 pairs from 378 attempts on 35 tasks derived from 2026 arXiv papers. The judge ordering agrees across the agent-attempt corpora (Spearman 0.62–0.73), asking judges for a rationale does not raise accuracy, and on the 16-judge leaderboard the newest frontier model leads at 71.0%, against a nine-judge majority of 67.3%.