TailGuard: Subgroup Tail Coverage Theory for Safe LLM Alignment
Charles Cao
Abstract
Safe preference alignment of large language models must handle rare harmful tails whose burden is concentrated on demographic, topical, linguistic, or adversarial subgroups. Existing coverage analyses operate on global density ratios that can look adequate on average while the harmful tail of a small group remains unsupported, leaving worst-group safety risk unidentified. We assume that the absolute harm scale is externally calibrated by severity labels or anchors and introduce TailGuard, a framework centered on subgroup quantile coverage $C_{g,\tau}$: the worst deployment-to-training density ratio restricted to group g's upper-τ harm tail. We prove that $C_{g,\tau}$ is a key statistical quantity governing subgroup tail safety: it governs tail-risk transfer, yields an offline impossibility result under missing tail support, and is necessary for fixed-policy subgroup tail-risk estimation. From a per-group CVaR_τ-constrained alignment objective, we derive a KL-regularized Gibbs form, motivate a Tail-Fisher active query rule that targets group-tail coverage deficits, and construct a per-group weighted CVaR risk-control deployment gate. We validate the mechanism through four completed data-level studies: synthetic coverage control, a semi-synthetic tail-hiding pilot on real safety-alignment labels, an observed-label active-query ablation, and an oracle-score risk-control-gate efficiency proxy. These experiments test the coverage mechanism and deployment diagnostic.
Chat is not available.
Successful Page Load