SIMPATIX: Transformer-Based Depression Prediction from Social Media Posts
Abstract
More than one in twenty adults worldwide are estimated to suffer from depression, affecting over 280 million people globally. Social media offers a unique opportunity to unobtrusively monitor psychological states, as individuals’ posts in the public domain can point to underlying and latent mental health issues. Understanding linguistic patterns in social media text may support early detection and intervention for depression. This study aimed to evaluate the effectiveness of a tuned transformer-based model in identifying depression indicators in social media posts. A manually curated dataset of several hundred anonymized posts was used. Self-declared depression patients as well as those without apparent indication of depression were selected for tuning the system. The model, based on RoBERTa, was fine-tuned on this dataset, with data augmentation applied to balance underrepresented classes. Model performance was assessed using accuracy, precision, recall, F1-score, ROC AUC, and PR AUC, and feature importance methods were applied to identify linguistic cues driving predictions. The fine-tuned model achieved 86% accuracy, with a ROC AUC of 0.88 and a PR AUC of 0.90, substantially outperforming the untuned model, which achieved 48% accuracy, and other large language models, which ranged from 41% to 46% accuracy. Feature importance analysis indicated that negative self-reference, affective language, and specific linguistic patterns were the most influential indicators of depressive behavior. These results demonstrate that domain-specific fine-tuning, data augmentation, and interpretability techniques enable robust detection of depression in social media text. Expanding dataset diversity and integrating multimodal signals, such as images and behavioral metadata, may further improve generalizability and support early mental health intervention strategies.