Burstiness Was Measured Wrong, and Prompting Cannot Aim It
Vittoria Lanzo ⋅ Dico Angelo
Abstract
Burstiness is a signal people already use to judge machine text, and its operational definition in practice is an unreviewed detector heuristic confounded with length: raw sentence-length variance grows with the mean. We treat sentence boundaries as a point process in word-time, scored with the Goh--Barab\'asi coefficient under the Kim--Jo finite-$n$ correction: bounded and scale-free, no model in the loop, screened by a Goodhart guard and a coherence gate. The protocol is paired evaluation on $200$ era-, topic-, and register-matched Wikinews articles dated 2016--2021: five prompting modes, nine model families, three seeds, uniform normalization of generated text before scoring. Scored raw, markdown scaffolding manufactures dispersion and produces a parity verdict that normalization reverses (instruction-mode $0.005$, $p=0.10$, becomes a significant $0.028$ deficit, $p=3\times10^{-5}$), and the uninstructed dispersion gap more than doubles ($0.043$ to $0.097$). Of the forty-five family-by-mode cells, thirty-two land significantly flatter than the human articles, eight show no significant difference, and five land significantly more dispersed: machine flatness is a strong central tendency that prompting moves without targeting. The same instruction leaves four families short, lands two on the human level (CI-bounded equivalence), moves one further away, pushes two past, and few-shot exemplars of typical human rhythm are the flattest mode by rank-biserial in seven of nine families. We model human rhythm; we do not build detector evasion.
Chat is not available.
Successful Page Load