HansardNER: Full-Coverage Benchmarking of Sinhala Parliamentary Named Entity Recognition
Abstract
We present HansardNER, a Sinhala parliamentary NER corpus of 4,000 speaker turns (1.49 million words, 48,161 mentions) with LLM-generated silver annotations over nine corpus-derived coarse classes, where Sinhala NER is still evaluated on four CoNLL classes. Parliamentary turns are long and XLM-R accepts at most 512 positions, so we pack complete sentences into label-blind windows of at most 510 content sub-tokens and reconstruct predictions over the complete turn. Every benchmark system is evaluated on the same 401 test turns and 4,197 test mentions with 100% word coverage. A full-coverage ladder from HMMs to finetuned encoders establishes XLM-R as the strongest backbone (73.96 micro-F1 with a linear softmax head against 67.40 for a feature CRF). A controlled threeseed comparison then tests linear, static-convolution and dynamic-convolution heads with softmax and CRF decoding under one fixed pipeline: all six heads lie between 73.96 and 74.65 micro-F1. Dynamic short convolution provides no statistically reliable improvement over linear or static heads (+0.04 vs. static softmax, 95% CI −0.70 to +0.82). Linear+CRF has the highest mean (+0.70 over linear softmax) but its interval includes zero (−0.27 to +1.70), and a FLERT-style turn-local context variant changes dynamic+CRF by +0.11 (−0.87 to +1.18). Performance measures agreement with the silver annotations rather than accuracy against adjudicated human gold labels. Corpus, code, windows and result files will accompany the camera-ready version.