Analyzing The Changes of AI Safety Language Across 4 Frontier Labs
Abstract
This paper traces how frontier AI labs have redefined and responded to AI risks across their published safety policy documents. In the absence of regulation or agreed-upon standards, major labs have practiced self-governance under the banner of safety. However, what counts as risk and how it should be managed have continued to shift. Growing regulatory and technical attention to safety does not, in itself, guarantee that these frameworks have moved toward greater safety. Existing work has largely compared frameworks across labs or cataloged their common and divergent features. Less attention has been paid to shifts in the strength of stated commitments or governance mechanisms, and to the policy and market conditions under which these changes occurred. To address this gap, we constructed a codebook centered on risk definitions, thresholds, and benchmarks, and applied it to 19 policy documents published between 2023 and 2026 by four US frontier AI labs (Anthropic, OpenAI, Google DeepMind, xAI). In practice, we find that the loosening or tightening of commitments in safety policies occurred in roughly equal measure. Labs also generally tightened future commitments while relaxing disclosure and evaluation rules. This paper examines how these lab-specific changes might align with their broader policy and market environment. Tracing how safety commitments quietly shift over time is as critical as evaluating what those policies initially set out. Ultimately, this study establishes an empirical foundation for understanding incremental policy revisions, informing both future regulatory oversight and public accountability as AI governance standards take shape.