Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?
Sourabrata Mukherjee, Atul Kr. Ojha, John Philip McCrae, Ondrej Dusek
Abstract
Text style transfer (TST) is the task of transforming a text to reflect a particular style while preserving its original content. Evaluating TSToutputs is a multidimensional challenge, requiring the assessment of style transfer accuracy, content preservation, and naturalness. Us-ing human evaluation is ideal but costly, as is common in other natural language processing (NLP) tasks; however, automatic metrics forTST have not received as much attention as metrics for, e.g., machine translation or summarization. In this paper, we examine both set ofexisting and novel metrics from broader NLP tasks for TST evaluation, focusing on two popular subtasks—sentiment transfer and detoxification—in a multilingual context comprising English, Hindi, and Bengali. By conducting meta-evaluation through correlation with hu-man judgments, we demonstrate the effectiveness of these metrics when used individually and in ensembles. Additionally, we investigatethe potential of large language models (LLMs) as tools for TST evaluation. Our findings highlight newly applied advanced NLP metrics andLLM-based evaluations provide better insights than existing TST metrics. Our oracle ensemble approaches show even more potential.- Anthology ID:
- 2025.naacl-srw.41
- Volume:
- Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop)
- Month:
- April
- Year:
- 2025
- Address:
- Albuquerque, USA
- Editors:
- Abteen Ebrahimi, Samar Haider, Emmy Liu, Sammar Haider, Maria Leonor Pacheco, Shira Wein
- Venues:
- NAACL | WS
- SIG:
- Publisher:
- Association for Computational Linguistics
- Note:
- Pages:
- 418–434
- Language:
- URL:
- https://preview.aclanthology.org/moar-dois/2025.naacl-srw.41/
- DOI:
- 10.18653/v1/2025.naacl-srw.41
- Cite (ACL):
- Sourabrata Mukherjee, Atul Kr. Ojha, John Philip McCrae, and Ondrej Dusek. 2025. Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics?. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 4: Student Research Workshop), pages 418–434, Albuquerque, USA. Association for Computational Linguistics.
- Cite (Informal):
- Evaluating Text Style Transfer Evaluation: Are There Any Reliable Metrics? (Mukherjee et al., NAACL 2025)
- PDF:
- https://preview.aclanthology.org/moar-dois/2025.naacl-srw.41.pdf