Measurable Writing Standards for AI-Native Venues
Abstract
Language models now draft manuscripts, assist review, and screen submissions, so the institutions of publishing are becoming AI-native. We argue that this transition inherits a communication problem that venues have not yet measured, and one that AI tooling does not correct. We analyze roughly 2.8 million arXiv papers (1991 to 2025), 30,595 NeurIPS papers (1987 to 2025), and 24.5 million PubMed abstracts (1990 to 2025) using classical readability formulas, nine writing style metrics, acronym density and reuse, citation counts, and an LLM as judge protocol. We report four findings. NeurIPS abstracts have become harder to read on every classical metric, with Flesch Reading Ease falling from approximately 24 in 1987 to approximately 8 in 2025 and sensational language rising by approximately 80 percent between 2015 and 2025. Acronym density rose from 0.33 to 4.15 per 100 words in titles and from 0.62 to 2.97 in abstracts, diverging from the non-ML baseline after 2016, and roughly 89 percent of ML acronyms are used fewer than ten times. More readable NeurIPS papers tend to be more cited, an association that survives a control for subfield. Crucially for an AI-native venue, the AI tools meant to help do not correct these trends. Six language models used as readability judges report that readability is flat or improving, in direct contradiction to every classical metric, and the divergence widens after 2022. We propose seven measurable, auditable standards that AI-native venues could pilot, and we argue that automated screening must be anchored in transparent human-grounded metrics rather than a language model’s own judgment.