Abstract
Gender-fair language, an evolving linguistic variation in German, fosters inclusion by addressing all genders or using neutral forms. However, there is a notable lack of resources to assess the impact of this language shift on language models (LMs) might not been trained on examples of this variation. Addressing this gap, we present Lou, the first dataset providing high-quality reformulations for German text classification covering seven tasks, like stance detection and toxicity classification. We evaluate 16 mono- and multi-lingual LMs and find substantial label flips, reduced prediction certainty, and significantly altered attention patterns. However, existing evaluations remain valid, as LM rankings are consistent across original and reformulated instances. Our study provides initial insights into the impact of gender-fair language on classification for German. However, these findings are likely transferable to other languages, as we found consistent patterns in multi-lingual and English LMs.- Anthology ID:
- 2024.emnlp-main.592
- Volume:
- Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing
- Month:
- November
- Year:
- 2024
- Address:
- Miami, Florida, USA
- Editors:
- Yaser Al-Onaizan, Mohit Bansal, Yun-Nung Chen
- Venue:
- EMNLP
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 10604–10624
- Language:
- URL:
- https://aclanthology.org/2024.emnlp-main.592
- DOI:
- 10.18653/v1/2024.emnlp-main.592
- Cite (ACL):
- Andreas Waldis, Joel Birrer, Anne Lauscher, and Iryna Gurevych. 2024. The Lou Dataset - Exploring the Impact of Gender-Fair Language in German Text Classification. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 10604–10624, Miami, Florida, USA. Association for Computational Linguistics.
- Cite (Informal):
- The Lou Dataset - Exploring the Impact of Gender-Fair Language in German Text Classification (Waldis et al., EMNLP 2024)
- PDF:
- https://preview.aclanthology.org/landing_page/2024.emnlp-main.592.pdf