Rubric-Align: Safety Alignment through Dynamically Co-Evolving Rubrics
Abstract
Recent progress in large language models (LLMs) has led to impressive performance across diverse domains. However, their deployment in critical areas such as healthcare, law, and education raises serious concerns regarding the potential generation of harmful content. While reinforcement learning (RL)-based safety alignment methods have been extensively studied, they face persistent challenges, including sparse reward signals, limited interpretability, and limited transferability across domains. In this paper, we introduce Rubric‑Align, a framework that replaces static binary scalar rewards with dynamic, natural‑language evaluation rubrics. Instead of assigning a single reward to each response, Rubric‑Align provides fine‑grained and interpretable feedback, thereby alleviating reward sparsity and improving the transparency of the alignment signal. A key feature of Rubric-Align is its dynamic rubric evolution mechanism. Rather than using fixed reward templates, Rubric-Align periodically refines prompt-specific rubrics based on the current policy behavior, keeping the supervision signal informative and aligned with the model’s evolving failure modes. Experiments on safety benchmarks and vertical-domain settings show that Rubric-Align improves robustness against harmful and jailbreak prompts while largely preserving general capabilities, suggesting that natural-language rubrics provide a transferable interface for adapting safety supervision to new safety subdomains.