Beyond the Word List: Context-Aware Measurement of Grievance Across Languages
Abstract
A word list cannot tell whether a post condemns violence or calls for it, yet word lists such as the Grievance Dictionary are how grievance, a warning sign in threat assessment, is measured at scale from online text. We show that their reported accuracy is partly an artefact of how they are tested. In an existing five-language pool of 2{,}000 items, the ``random'' half turns out to be exactly the text the lexicon does not match, so the lexicon's macro-AUROC of 0.686 on the matched half falls to 0.500 on the other: a floor fixed by construction, not by language. We keep the dictionary's 22 constructs but replace term matching with models that read the target sentence inside its full post, with one multilingual encoder and one label space across English, Dutch, German, Italian and French, and we evaluate on a new benchmark whose test items the lexicon did not choose. Context helps most where the lexicon is silent: average precision on lexicon-negative text rises from 0.14 to 0.20, with the largest gains on quoted, implicit and cross-sentence grievance, the cases a word list cannot see. In two protest streams, Australian English and Indonesian, a language with no validated lexicon, only the contextual measure moves with offline mobilisation. Grievance is measured more faithfully by reading context, and more honestly on text the lexicon did not select. Code and benchmark: \mbox{\url{https://anonymous.4open.science/r/multilingual_grievance-4564/}}