How Should the Watchmen be Watched? Rethinking the Role of Humans in Automatic Evaluation of LLMs
Abstract
Automatic evaluation has been central to the rapid development of language models, but its validity still relies heavily on annotation-based human meta-evaluation. This is problematic because benchmarks now saturate faster and increasingly target open-ended tasks such as tutoring and mental health support, where annotation scales poorly and informs little. This position paper argues that human effort in meta-evaluation should be allocated according to two task dimensions: open-endedness, how loosely a task specifies its valid outputs and goals, and verifiability, how objectively an output's correctness can be established. Open-endedness enlarges the space annotators must cover, and unverifiability makes each annotation less reliable. Where both are high, we argue that effort is better invested in construct validity and extrinsic meta-evaluation against real-world data and outcomes than in annotation. Yet, among 1,294 papers from top ML venues that perform meta-evaluation, we find no evidence that extrinsic meta-evaluation is more prevalent for open-ended or unverifiable tasks. Using mental health support as a case study, we illustrate how benchmark and meta-evaluation design can be adapted to open-endedness and verifiability, and propose practical guidelines for allocating human effort accordingly.