SafetyTrace: Mapping Safety Behavior and Reversal Effects Across Ordered LLM Interventions
Abstract
Conversational large language models are often evaluated under a fixed safety prompt or guardrail configuration, even though deployed systems may combine and tune several safeguards. We study how model behavior changes as those safe- guards become progressively more restrictive. SafetyTrace represents each prompt as a trajectory across ordered system-prompt interventions and defines reversal magnitude as the largest later increase in disallowed compliance. We first identify high-reversal cases in an exploratory screen, then evaluate a fixed set of cases again using fresh generations and a 3-of-4 consensus across four automated judges. The replicated trajectories closely match the discovery trajectories, with mean Pearson correlation r = 0.920. All six replicated cases reject a monotone non-increasing model after Benjamini-Hochberg correction, with the largest adjusted q = 0.00120. In a separate experiment where each intervention retains the restrictions introduced at earlier levels, two cases still show 0.50 reversals at q = 0.045985. These results show that progressively stronger safety instructions do not always produce a smooth reduction in disallowed compliance, and that evaluating only a single configuration can miss failures that appear at intermediate intervention levels.