The Coverage Selection Gap: Ranking, Not Generation, Limits LLM Rubric Construction
Abstract
Task-specific rubrics are increasingly used to evaluate open-ended language-model outputs, yet constructing high-coverage rubrics from complex source documents remains largely manual. We study automatic rubric construction in professional M&A due diligence, where a system must identify criteria for evaluating a deliverable without access to an expert-authored reference rubric. We introduce a multi-axis LLM framework that searches complementary structural and free-form views of the source documents to construct a broad candidate pool, followed by reference-free ranking. Developed on 16 tasks and evaluated on four held-out tasks from a 20-task benchmark, the resulting candidate pools achieve 0.876 macro recall against senior-authored reference rubrics under our LLM-based criterion-equivalence protocol. However, the covered reference criteria are poorly concentrated near the top of the reference-free ranking. At a selection budget of k = 60, our best reference-free ranker achieves 0.349 recall, whereas a greedy reference-informed ordering of the same candidate pools reaches their measured 0.876 recall ceiling. This 52.7-percentage-point oracle gap quantifies coverage present in the pools but not surfaced by the reference-free ranker. In our ablations, adding individual structural search axes improves recall by 6–10 percentage points, while increasing candidate volume along existing search directions yields at most 1.5 points and can reduce F1. On the two tasks for which complete outputs from both systems are available, the union of our candidates and those from a single-call GPT baseline achieves 0.934 recall. These findings indicate that, in the evaluated document-grounded M&A setting, candidate selection is a substantial bottleneck once a high-coverage pool has been constructed.