Alignment Is Relational: Auditing Safety Guards Across Human Constituencies
Sagnik Bhattacharya
Abstract
Open guard models mediate what deployed language-model systems treat as "harmful," emitting a single safe/unsafe verdict, yet human judgments diverge sharply: on DICES-350, racial group majorities directly oppose each other on 77 of 350 conversations (22\%). We audit four open guards against disaggregated rather than collapsed human labels and find three patterns. First, every guard shows substantial alignment differences across racial-group majority targets (full-set AUC gaps .03--.14, widening to .14--.23 on group-conflicted items; all bootstrap CIs exclude zero). Second, guard confidence tracks rater disagreement in guard-specific ways, from WildGuard ($\rho = -.44$) to ShieldGemma-9B ($\rho = +.12$, uncorrected $p = .02$). Third, apparent differences in which group a guard best tracks are descriptive only (non-significant at $K=5$) and should not be read as distinct ``bias profiles.'' An achievable-alignment analysis shows DICES-350's disagreement structure does not itself preclude substantially more balanced binary agreement, giving a model-free reference for interpreting guard disparities. We release the audit harness and argue guard model cards should report perspectival alignment profiles alongside aggregate performance.
Chat is not available.
Successful Page Load