Claims Without Artefacts: Tiered Disclosure for Frontier AI Safety Results
Abstract
Twelve leading AI companies published or updated frontier AI safety frameworks in 2025. Over the same period the Foundation Model Transparency Index recorded a sector average of 40/100, six major developers at zero on model-information disclosures, and no developer adequately reporting train-test overlap. The claims that carry the most weight in AI safety are the ones a reader can least often rerun. We call this the evidential inversion, and we treat it as an evaluation-methodology failure with technical evidence behind it: attack-success-rate comparisons across systems are often founded on low-validity measurements, frontier models can defeat every output-based jailbreak monitor tested, and the 2026 International AI Safety Report concludes that reliable pre-deployment testing has become harder to conduct. We turn the usual call for disclosure into something a programme committee can actually run. Authors of papers making such claims file a claim inventory, a scope statement, and a declared disclosure tier: public, controlled, or claim-restricted. A gating checklist question keeps the cost near zero for submissions the standard does not touch. Controlled review is hosted by a federated colloquium of existing secure-review entities under a bounded pilot. Each requirement is stated with what it costs per submission, who performs the check, and what happens when it fails, and we name the items the proposal does not yet cost out.