Calibration Without Discrimination:large Instance-Level Uncertainty Failures in LLMs Across Global South Constitutions
DEBJYOTI PAUL
Abstract
Large language models are rapidly becoming infrastructure for everyday information access across Global South jurisdictions. Unlike scientific or factual knowledge, fundamental constitutional principles are not globally uniform: the rights guaranteed under India's Art.21, China's PRC Constitution Arts. 33-56, Brazil's Art. 5, and the EU Charter are jurisdiction-specific, locally grounded, and cannot be inferred from general capability scores. Whether LLMs are calibrated to each jurisdiction's constitutional framework-and whether they know the limits of their own knowledge within it-is a governance-level question that no benchmark currently measures. We introduce \textbf{ConstitutionBench}: 9,000 scenarios across five jurisdictions (US, India, China, Brazil, EU), producing 65,997 labeled responses from three model families (Llama-3.3-70B, Claude Haiku~4.5, GPT-4o-mini) plus jurisdiction-matched specialists (Sarvam-105B, Qwen-2.5-72B). We propose two metrics: \textbf{CCE} (Constitutional Calibration Error, $\text{CCE}(j) = h_j - \hat{h}_j$, where $h_j$ is the hedge rate and $\hat{h}_j$ the non-recognized citation rate) and $\rho_j = P(\text{hedge} \mid \text{non-recognized}) - P(\text{hedge} \mid \text{recognized})$ (instance-level calibration: whether hedges fall on the right scenarios). We find a striking \textbf{dissociation}: China (worst aggregate $\text{CCE}$) is the \emph{only} instance-calibrated jurisdiction ($\rho_j > 0$); US, EU, and Brazil are instance-anti-calibrated ($\rho_j < 0$) despite acceptable aggregate scores. Llama-3.3-70B's CCE range is $8.8\times$ greater than GPT-4o-mini's despite higher MMLU scores---capability benchmarks do not predict calibration. All five jurisdictions fail our governance rubric on distinct axes.
Chat is not available.
Successful Page Load