Emily Sofi Ohman
2026
Quality and Agreement in Multilabel Emotion Annotation: A Case Study and Evaluation Framework
Emily Sofi Ohman | Anna Koufakou
Proceedings of Computational Affective Science (CAS) @ LREC 2026
Emily Sofi Ohman | Anna Koufakou
Proceedings of Computational Affective Science (CAS) @ LREC 2026
Emotion annotation is inherently subjective, yet most NLP pipelines still assume “gold” labels, typically produced by majority voting, and treat annotator variation as noise. In this paper, we present a multilabel emotion annotation case study and use it to examine how annotator behavior and aggregation choices affect both agreement estimates and downstream emotion classifiers. Rather than collapsing disagreement into a single label, we represent targets as soft vote-share labels (including an intensity-weighted variant) and evaluate models using both thresholded metrics (macro-/micro-F1) and probabilistic alignment (Bernoulli cross-entropy SoftBCE), alongside data-derived disagreement diagnostics. Across annotation regimes, we show that disagreement is structured and leaves measurable traces in model behavior: hard labels may maximize F1 metrics, while soft supervision yields predictions that better reflect empirical annotator variance and uncertainty. Our results provide practical guidance for designing, aggregating, and evaluating multilabel emotion datasets when multiple interpretations are plausible.
2024
Text Length and the Function of Intentionality: A Case Study of Contrastive Subreddits
Emily Sofi Ohman | Aatu Liimatta
Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities
Emily Sofi Ohman | Aatu Liimatta
Proceedings of the 4th International Conference on Natural Language Processing for Digital Humanities
Text length is of central concern in natural language processing (NLP) tasks, yet it is very much under-researched. In this paper, we use social media data, specifically Reddit, to explore the function of text length and intentionality by contrasting subreddits of the same topic where one is considered more serious/professional/academic and the other more relaxed/beginner/layperson. We hypothesize that word choices are more deliberate and intentional in the more in-depth and professional subreddits with texts subsequently becoming longer as a function of this intentionality. We argue that this has deep implications for many applied NLP tasks such as emotion and sentiment analysis, fake news and disinformation detection, and other modeling tasks focused on social media and similar platforms where users interact with each other via the medium of text.