Trustworthy Ground Truth? The Limits of Human Validation in LLM-Based Legal Metadata Extraction in a Resource-Constrained Local Deployment
Abstract
The adoption of Large Language Models (LLMs) to populate structured metadata in institutional legal corpora is a low-cost strategy, particularly relevant in Global South contexts where access to large-scale professional annotation and commercial extraction tools may be constrained. However, the validity of these systems is typically evaluated against a human "ground truth" whose own reliability is rarely audited. We present an inter-annotator agreement analysis of eight legal fields produced by a locally deployed LLM-based extraction pipeline (qwen2.5:14b) from 1,087 court rulings issued by a Latin American state institution. Two non-expert annotators (interns) independently validated an overlapping set of 653 rulings. We find that Cohen's kappa (0.324) yields a much lower value than the observed agreement under the high prevalence of a single category—a clear instance of the "kappa paradox"—whereas Gwet's AC1 (0.778) is more consistent with the raw agreement rate. More importantly, in fields requiring legal interpretation or list-completeness judgments, unresolved human disagreement is substantially more frequent than consensus that the extracted value is invalid (grounds for the claim: 41.5%; panel judges: 43.8% human disagreement), whereas simple entity-extraction fields show high reliability (reporting judge: 6.0% human disagreement). We argue that, in resource-constrained contexts, the quality of the human validation process is a critical component of the reliability of deployed legal AI systems, alongside model performance.