Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits

Bohan Li; Jiannan Guan; Longxu Dou; Yunlong Feng; Dingzirui Wang; Yang Xu; Enbo Wang; Qiguang Chen; Bichen Wang; Xiao Xu; Yimeng Zhang; Libo Qin; Yanyan Zhao; Qingfu Zhu; Wanxiang Che (车万翔)

Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits

Bohan Li, Jiannan Guan, Longxu Dou, Yunlong Feng, Dingzirui Wang, Yang Xu, Enbo Wang, Qiguang Chen, Bichen Wang, Xiao Xu, Yimeng Zhang, Libo Qin, Yanyan Zhao, Qingfu Zhu, Wanxiang Che

Abstract

The Myers-Briggs Type Indicator (MBTI) is one of the most influential personality theories reflecting individual differences in thinking, feeling, and behaving. MBTI personality detection has garnered considerable research interest and has evolved significantly over the years. However, this task tends to be overly optimistic, as it currently does not align well with the natural distribution of population personality traits. Specifically, the self-reported labels in existing datasets result in data quality issues and the hard labels fail to capture the full range of population personality distributions. In this paper, we identify the task by constructing MBTIBench, the first manually annotated MBTI personality detection dataset with soft labels, under the guidance of psychologists. Our experimental results confirm that soft labels can provide more benefits to other psychological tasks than hard labels. We highlight the polarized predictions and biases in LLMs as key directions for future research.

Anthology ID:: 2025.coling-main.339
Volume:: Proceedings of the 31st International Conference on Computational Linguistics
Month:: January
Year:: 2025
Address:: Abu Dhabi, UAE
Editors:: Owen Rambow, Leo Wanner, Marianna Apidianaki, Hend Al-Khalifa, Barbara Di Eugenio, Steven Schockaert
Venue:: COLING
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 5071–5081
Language:
URL:: https://preview.aclanthology.org/remove-affiliations/2025.coling-main.339/
DOI:
Bibkey:
Cite (ACL):: Bohan Li, Jiannan Guan, Longxu Dou, Yunlong Feng, Dingzirui Wang, Yang Xu, Enbo Wang, Qiguang Chen, Bichen Wang, Xiao Xu, Yimeng Zhang, Libo Qin, Yanyan Zhao, Qingfu Zhu, and Wanxiang Che. 2025. Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits. In Proceedings of the 31st International Conference on Computational Linguistics, pages 5071–5081, Abu Dhabi, UAE. Association for Computational Linguistics.
Cite (Informal):: Can Large Language Models Understand You Better? An MBTI Personality Detection Dataset Aligned with Population Traits (Li et al., COLING 2025)
Copy Citation:
PDF:: https://preview.aclanthology.org/remove-affiliations/2025.coling-main.339.pdf

PDF Search Fix data