Geometric Trust: Hidden-State Geometry for Post-Deployment Trust Estimation in Instruction-Tuned LLM Classifiers
Abstract
Instruction-tuned large language models (LLMs) are increasingly used as classifiers in high-stakes domains, yet classification accuracy alone does not indicate whether an individual prediction can be trusted. Existing confidence measures, such as token probabilities and perplexity, rely primarily on model outputs and do not fully exploit the information encoded in hidden representations. We present Geometric Trust, a training-free framework for post-deployment trust estimation in instruction-tuned LLM classifiers. We systematically compare probability-based and hidden-state trust measures, including Decision Margin, Feature Distance to the Decision Boundary (fDBD), and the Normalized Cosine Index (NCI), and introduce Topic Margin Trust, a document-level confidence measure based on the separation between the two highest-ranked candidate topics. Experiments on a newly constructed benchmark for fine-grained Sustainability Accounting Standards Board (SASB) ESG disclosure classification show that Decision Margin provides the strongest overall trust signal, while fDBD is the strongest hidden-state geometric measure and remains competitive with probability-based confidence. Topic Margin Trust further improves document-level confidence estimation over conventional ranking scores. Finally, the proposed trust measures enable effective selective prediction, reducing the classification error at 90% coverage from 1.72% to 0.19% for Llama-3.1 and from 1.94% to 0.13% for Qwen2.5. We will publicly release the benchmark, evaluation framework, and implementation to support future research on trustworthy LLM classification.