Saad Mankarious
2026
MindSET: Advancing Mental Health Benchmarking through Large-Scale Social Media Data
Saad Mankarious | Edward Kempa | Daniel Wiechmann | Elma Kerz | Yu Qiao | Ayah Zirikly
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Saad Mankarious | Edward Kempa | Daniel Wiechmann | Elma Kerz | Yu Qiao | Ayah Zirikly
Proceedings of the Fifteenth Language Resources and Evaluation Conference
Social media data has become a vital resource for studying mental health, offering real-time insights into thoughts, emotions, and behaviors that traditional methods often miss. Progress in this area has been facilitated by benchmark datasets for mental health analysis; however, most existing benchmarks have become outdated due to limited data availability, inadequate cleaning, and the inherently diverse nature of social media content (e.g., multilingual and harmful material). We present a new benchmark dataset, MindSET, curated from Reddit using self-reported diagnoses to address these limitations. The annotated dataset contains over 13M annotated posts across seven mental health conditions—more than twice the size of previous benchmarks. To ensure data quality, we applied rigorous preprocessing steps, including language filtering, and removal of Not Safe for Work (NSFW) and duplicate content. We further performed a linguistic analysis using LIWC to examine psychological term frequencies across the eight groups represented in the dataset. To demonstrate the dataset’s utility, we conducted binary classification experiments for diagnosis detection using both fine-tuned language models and Bag-of-Words (BoW) features. Models trained on MindSET consistently outperformed those trained on previous benchmarks, achieving up to an 18-point improvement in F1 for Autism detection. Overall, MindSET provides a robust foundation for researchers exploring the intersection of social media and mental health, supporting both early risk detection and deeper analysis of emerging psychological trends.
Mirroring Minds: Asymmetric Linguistic Accommodation and Diagnostic Identity in ADHD and Autism Reddit Communities
Saad Mankarious | Nour Zeid | Iyad Ait Hou | Rebecca Hwa | Ayah Zirikly
Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology (CLPsych 2026)
Saad Mankarious | Nour Zeid | Iyad Ait Hou | Rebecca Hwa | Ayah Zirikly
Proceedings of the 11th Workshop on Computational Linguistics and Clinical Psychology (CLPsych 2026)
Social media research on mental health has focused predominantly on detecting and diagnosing conditions at the individual level. In this work, we shift attention to intergroup behavior, examining how two prominent neurodivergent communities, ADHD and autism, adjust their language when engaging with each other on Reddit. Grounded in Communication Accommodation Theory (CAT), we first establish that each community maintains a distinct linguistic profile as measured by the Linguistic Inquiry and Word Count (LIWC) dictionary. We then show that these profiles shift in opposite directions when users cross community boundaries: features that are elevated in one group’s home community decrease when its members post in the other group’s space, and vice versa, consistent with convergent accommodation. Finally, in an exploratory longitudinal analysis around the moment of public diagnosis disclosure, we find that its effects on linguistic style are small and, in some cases, directionally opposite to cross-community accommodation, providing initial evidence that situational audience adaptation and longer-term identity processes may involve different mechanisms. Our findings contribute to understanding intergroup communication dynamics among neurodivergent populations online and carry implications for community moderation and clinical perspectives on these conditions.