TrustMod-SM: A Multi-Axis Benchmark for Evaluating Trustworthiness of LLMs in Social Media Content Moderation
Abstract
Large language models have shown remarkable promise for social media content moderation, with recent studies demonstrating performance rivaling human annotators on policy-compliance tasks. However, the trustworthiness of these models beyond aggregate accuracy remains largely unexamined, posing significant risks, a model with high overall accuracy may still exhibit substantial error-rate disparities across demographic groups while remaining vulnerable to adversarial manipulation. In this paper, we introduce TrustMod-SM and aim to comprehensively evaluate the trustworthiness of LLM-based content moderators across five dimensions: trustfulness, fairness, safety, robustness, and context integrity. TrustMod-SM comprises about 29K evaluation instances curated from eight established datasets, and we evaluate thirteen open-weight models (0.5B-14B parameters), including both text-only and vision-language architectures. Our analysis reveals that models consistently exhibit trustworthiness concerns, often displaying greater demographic disparity as detection accuracy improves, complying with both concealment and exaggeration inducements, assigning high confidence to incorrect predictions while rarely escalating ambiguous content for human review, and failing to distinguish hate from counter-speech and reclaimed language, particularly at sub-8B scales. Furthermore, scaling does not uniformly improve trustworthiness, and a cross-dimensional analysis reveals non-trivial associations between axes, challenging the independence assumed by prior benchmarks. We publicly release our benchmark data, evaluation code, and model predictions at https://anonymous.4open.science/r/TrustMod_SM WARNING: This paper contains model outputs that may be considered offensive.