Systematic Literature Review for Computational Detection of Anthropomorphic Language
Abstract
Anthropomorphic language the attribution of human characteristics to non-human entities is increasingly prevalent in how large language models (LLMs) are described, deployed, and perceived. However, despite being extensively researched by psychologists for decades, computational detection of anthropomorphism in natural language is still an emerging research domain. Anthropo-morphic framing can be expressed through linguistic indicators such as the use of certain pronouns, verbs and sentence constructions; therefore, NLP provides a good basis for the development of systematic detection techniques. This paper offers a systematic review of the scientific literature on the detection of anthropomorphism through computation and proposes a classification of available methods into five methodological paradigms: crowdsourcing and classification, unsupervised detection, automated metric-based detection, supervised classifier-based detection, and taxonomy and intervention frameworks. The review reveals four major open problem, namely: the lack of validated detection of anthropomorphism in conversational text generated by LLMs, the absence of methods of cross-lingual detection, the lack of a large-scale multi-domain benchmark, and the lack of empirical connection between detection accuracy and reduced downstream harm (including over trust and emotional dependence). These open problems are particularly crucial, taking into account the scale at which LLMs are currently used among diverse populations and in many important domains.